Dergiler / Turkish Journal of Electrical Engineering and Computer Sciences / 2017 / Cilt: 25 - Sayı: 2

Hadoop framework implementation and performance analysis on a cloud

Sayfa
705–716
DOI
—

Abstract

The Hadoop framework uses the MapReduce programming paradigm to process big data by distributingdata across a cluster and aggregating. MapReduce is one of the methods used to process big data hosted on largeclusters. In this method, jobs are processed by dividing into small pieces and distributing over nodes. Parameterssuch as distributing method over nodes, the number of jobs held in a parallel fashion, and the number of nodes in thecluster affect the execution time of jobs. The aim of this paper is to determine how the numbers of nodes, maps, andreduces affect the performance of the Hadoop framework in a cloud environment. For this purpose, tests were carriedout on a Hadoop cluster with 10 nodes hosted in a cloud environment by running PiEstimator, Grep, Teragen, andTerasort benchmarking tools on it. These benchmarking tools available under the Hadoop framework are classi ed asCPU-intensive and CPU-light applications as a result of tests. In CPU-light applications, increasing the numbers ofnodes, maps, and reduces does not improve the efficiency of these applications; they even cause an increase in time spenton jobs by using system resources unnecessarily. Therefore, in CPU-light applications, selecting the numbers of nodes,maps, and reduces as minimum is found as the optimization of time spent on a process. In CPU-intensive applications,according to the phase that small job pieces is processed, it is found that selecting the number of maps or reduces equalto total number of CPUs on a cluster is the optimization of time spent on a process.