Communication-Efficient Cluster Scalable Genomics Data Processing Using Apache Arrow Flight

Ahmad, T.; Ma, Chengxin; Al-Ars, Zaid; Peter Hofstee, H. Peter

doi:10.1109/ISPDC55340.2022.00028

Communication-Efficient Cluster Scalable Genomics Data Processing Using Apache Arrow Flight

Conference paper (2022)

Authors

T. Ahmad Computer Engineering

Chengxin Ma Student

Zaid Al-Ars Computer Engineering

H. Peter Peter Hofstee Computer Engineering

Research Group

Computer Engineering

DOI: https://doi.org/10.1109/ISPDC55340.2022.00028

Big Data Parallel Processing Genomics Apache Arrow Whole Genome/Exome Sequencing In-Memory Plasma Object Store

To reference this document use:

http://resolver.tudelft.nl/uuid:867e7f25-6b2c-4769-8018-8f5ad2d2ac38

More Info

expand_more

Published Date

2022

Language

English

Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Research Group

Computer Engineering

Abstract

Current cluster scaled genomics data processing solutions rely on big data frameworks like Apache Spark, Hadoop and HDFS for data scheduling, processing and storage. These frameworks come with additional computation and memory overheads by default. It has been observed that scaling genomics dataset processing beyond 32 nodes is not efficient on such frameworks.To overcome the inefficiencies of big data frameworks for processing genomics data on clusters, we introduce a low-overhead and highly scalable solution on a SLURM based HPC batch system. This solution uses Apache Arrow as in-memory columnar data format to store genomics data efficiently and Arrow Flight as a network protocol to move and schedule this data across the HPC nodes with low communication overhead.As a use case, we use NGS short reads DNA sequencing data for pre-processing and variant calling applications. This solution outperforms existing Apache Spark based big data solutions in term of both computation time (2x) and lower communication overhead (more than 20-60% depending on cluster size). Our solution has similar performance to MPI-based HPC solutions, with the added advantage of easy programmability and transparent big data scalability. The whole solution is Python and shell script based, which makes it flexible to update and integrate alternative variant callers. Our solution is publicly available on GitHub at https://github.com/abs-tudelft/time-to-fly-high/tree/main/genomics

Files

Communication_Efficient_Cluste... (pdf)

(pdf | 1.24 Mb)

- Embargo expired in 01-07-2023

Unknown license