Journals / Turkish Journal of Electrical Engineering and Computer Sciences / 2018 / Cilt: 26 - Sayı: 5
Optimization in the catalyst optimizer of Spark SQL
- Pages
- 2489–2499
- DOI
- —
Abstract
Apache Spark is one of the most technically challenged frameworks for cluster computing in which dataare processed in a parallel fashion. The cluster consists of unreliable machines. It processes a large amount of datafaster compared to the MapReduce framework. For providing the facility of optimized and fast SQL query processing,a new unit is developed in Apache Spark named Spark SQL. It allows users to use relational processing and functionalprogramming in one place. It provides many optimizations by leveraging the benefits of its core. This is called thecatalyst optimizer. This optimizer has many rules to optimize queries for efficient execution. In this paper, we discussa scenario in which the catalyst optimizer is not able to optimize the query competently for a specific case. This is thereason for inefficient memory usage and increases in the time required for the execution of the query by Spark SQL. Fordealing with this issue, we propose a solution in this paper by which the query is optimized up to the peak level. Thissignificantly reduces the time and memory consumed by the shuffling process