Toward Efficient In-memory Data Analytics on NUMA Systems

Memarzia, Puya; Ray, Suprio; Bhavsar, Virendra C

Computer Science > Databases

arXiv:1908.01860 (cs)

[Submitted on 5 Aug 2019 (v1), last revised 25 Jan 2020 (this version, v3)]

Title:Toward Efficient In-memory Data Analytics on NUMA Systems

Authors:Puya Memarzia, Suprio Ray, Virendra C Bhavsar

View PDF

Abstract:Data analytics systems commonly utilize in-memory query processing techniques to achieve better throughput and lower latency. Modern computers increasingly rely on Non-Uniform Memory Access (NUMA) architectures in order to achieve scalability. A key drawback of NUMA architectures is that many existing software solutions are not aware of the underlying NUMA topology and thus do not take full advantage of the hardware. Modern operating systems are designed to provide basic support for NUMA systems. However, default system configurations are typically sub-optimal for large data analytics applications. Additionally, achieving NUMA-awareness by rewriting the application from the ground up is not always feasible.
In this work, we evaluate a variety of strategies that aim to accelerate memory-intensive data analytics workloads on NUMA systems. We analyze the impact of different memory allocators, memory placement strategies, thread placement, and kernel-level load balancing and memory management mechanisms. With extensive experimental evaluation we demonstrate that methodical application of these techniques can be used to obtain significant speedups in four commonplace in-memory data analytics workloads, on three different hardware architectures. Furthermore, we show that these strategies can speed up two popular database systems running a TPC-H workload.

Comments:	15 pages, 9 figures Rev 3: fixed minor typos
Subjects:	Databases (cs.DB)
ACM classes:	H.2.4
Cite as:	arXiv:1908.01860 [cs.DB]
	(or arXiv:1908.01860v3 [cs.DB] for this version)
	https://doi.org/10.48550/arXiv.1908.01860

Submission history

From: Puya Memarzia [view email]
[v1] Mon, 5 Aug 2019 21:03:00 UTC (2,584 KB)
[v2] Wed, 7 Aug 2019 01:34:48 UTC (2,572 KB)
[v3] Sat, 25 Jan 2020 06:05:31 UTC (2,571 KB)

Computer Science > Databases

Title:Toward Efficient In-memory Data Analytics on NUMA Systems

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Databases

Title:Toward Efficient In-memory Data Analytics on NUMA Systems

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators