Improving Spark Application Throughput Via Memory Aware Task Co-location - Research Portal

Associated organisational units

Electronic data

middleware17-paper74
Rights statement: © Owner/Author ACM, 2017. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Middleware '17 Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference http://dx.doi.org/10.1145/3135974.3135984
Accepted author manuscript, 14 MB, PDF document
Available under license: CC BY-NC: Creative Commons Attribution-NonCommercial 4.0 International License

Text available via DOI:

https://doi.org/10.1145/3135974.3135984
Final published version

View graph of relations

Improving Spark Application Throughput Via Memory Aware Task Co-location: A Mixture of Experts Approach

Research output: Contribution in Book/Report/Proceedings - With ISBN/ISSN › Conference contribution/Paper › peer-review

Published

More...

Publication date	11/12/2017
Host publication	Middleware '17 Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference
Place of Publication	New York
Publisher	ACM
Pages	95-108
Number of pages	14
ISBN (print)	9781450347204
<mark>Original language</mark>	English

Abstract

Data analytic applications built upon big data processing frameworks such as Apache Spark are an important class of applications. Many of these applications are not latency-sensitive and thus can run as batch jobs in data centers. By running multiple applications on a computing host, task co-location can significantly improve the server utilization and system throughput. However, effective task co-location is a non-trivial task, as it requires an understanding of the computing resource requirement of the co-running applications, in order to determine what tasks, and how many of them, can be co-located. State-of-the-art co-location schemes either require the user to supply the resource demands which are often far beyond what is needed; or use a one-size-fits-all function to estimate the requirement, which, unfortunately, is unlikely to capture the diverse behaviors of applications.

In this paper, we present a mixture-of-experts approach to model the memory behavior of Spark applications. We achieve this by learning, off-line, a range of specialized memory models on a range of typical applications; we then determine at runtime which of the memory models, or experts, best describes the memory behavior of the target application. We show that by accurately estimating the resource level that is needed, a co-location scheme can effectively determine how many applications can be co-located on the same host to improve the system throughput, by taking into consideration the memory and CPU requirements of co-running application tasks. Our technique is applied to a set of representative data analytic applications built upon the Apache Spark framework. We evaluated our approach for system throughput and average normalized turnaround time on a multi-core cluster. Our approach achieves over 83.9% of the performance delivered using an ideal memory predictor. We obtain, on average, 8.69x improvement on system throughput and a 49% reduction on turnaround time over executing application tasks in isolation, which translates to a 1.28x and 1.68x improvement over a state-of-the-art co-location scheme for system throughput and turnaround time respectively.

Bibliographic note

© Owner/Author ACM, 2017. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Middleware '17 Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference http://dx.doi.org/10.1145/3135974.3135984

Research

Associated organisational units

Electronic data

Links

Text available via DOI:

Improving Spark Application Throughput Via Memory Aware Task Co-location: A Mixture of Experts Approach

Abstract

Bibliographic note

Quick Links

Connect With Us

Faculties & Depts

Contact Us