Characterizing the Effectiveness of Query Optimizer in Spark

2018 IEEE World Congress on Services (SERVICES)(2018)

引用 2|浏览31
暂无评分
摘要
In the big data community, Spark has been widely used for processing interactive queries. Spark employs a query optimizer, called Catalyst, to provides a set of optimization rules and supports Cost-Based Optimization (CBO). In this paper, we investigated the effectiveness of the optimization rules and cost-based optimization in Catalyst. We conducted comprehensive validation experiments by varying the data volume and cluster scale, and found that the execution time of most TPC-H queries were reduced slightly even when query optimizations are applied. We derived some interesting observations on Catalyst, which can help the community better understand and improve the query optimizer of Spark in future.
更多
查看译文
关键词
spark,catalyst,query optimization
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要