Статья

Big Data Analytics in the Cloud: Spark on Hadoop vs MPI/OpenMP on Beowulf

Jorge Luis Reyes-OrtizDIBRIS - University of Genoa, Via Opera Pia 13, I-16145 Genoa, ItalyLuca OnetoDITEN - University of Genoa, Via Opera Pia 11A, I-16145 Genoa, ItalyDavide AnguitaDIBRIS - University of Genoa, Via Opera Pia 13, I-16145 Genoa, Italy

2015en

ABI

Аннотация

One of the biggest challenges of the current big data landscape is our inability to process vast amounts of information in a reasonable time. In this work, we explore and compare two distributed computing frameworks implemented on commodity cluster architectures: MPI/OpenMP on Beowulf that is high-performance oriented and exploits multi-machine/multicore infrastructures, and Apache Spark on Hadoop which targets iterative algorithms through in-memory computing. We use the Google Cloud Platform service to create virtual machine clusters, run the frameworks, and evaluate two supervised machine learning algorithms: KNN and Pegasos SVM. Results obtained from experiments with a particle physics data set show MPI/OpenMP outperforms Spark by more than one order of magnitude in terms of processing speed and provides more consistent performance. However, Spark shows better data management infrastructure and the possibility of dealing with other aspects such as node failure and data replication.

Перевод пока недоступен

Идентификаторы

DOI: 10.1016/j.procs.2015.07.286

Цитирования и источники

Цитирований: 2Использованных источников: 0

Показатели — AkademScholar