AI Briefing
KO

Kakao Data PE Cell Publishes Spark Shuffle Partition Optimization Experiment Results

·2021.10.08 00:00

Key point

Kakao Data PE Cell shared a case study where query improvements and partition count adjustments reduced execution time from 8.4 hours to 3.5 hours through Spark Shuffle Partition optimization experiments.

1 / 12

Details

Logan from the Kakao Data PE Cell (Applied Analytics Team) shared a strategy to prevent Shuffle Spill by adjusting partition count and size, which are core aspects of Spark resource configuration. In an actual HDFS cluster environment (3 Cores, 6GB Memory), a 4-stage experiment was conducted on a join query between two tables. With the initial setting (300 partitions), 770GB of Shuffle Spill occurred, taking 8.4 hours. Increasing the partition count to 1800 and reducing the size to 140MB decreased Spill to 220GB and shortened execution time to 5.5 hours. Applying query optimization (GroupBy aggregation before Join) eliminated Spill and further reduced execution time to 3.5 hours. After optimization, similar performance of 3.7 hours was maintained even when memory per Core was reduced to 3GB. The author suggested an optimization priority order of 'Query > Partition count adjustment > Memory per Core increase' and recommended maintaining Shuffle Partition size between 100~200MB.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.