Kurly Cuts Data Pipeline Processing Time to 13 Seconds with BigQuery Adoption
Key point
Kurly reduced data loading time from 30 minutes to 13 seconds and improved cost efficiency by adopting BigQuery.
Details
Kurly redesigned its data pipeline architecture by adopting BigQuery, achieving major improvements in performance and cost. The company built separate pipelines for structured and unstructured data, and uses AWS DMS and Kafka to collect change logs.
The existing Data Warehouse's UPSERT approach was script-based, resulting in slow processing speeds and data synchronization delays. In contrast, by using BigQuery's Merge syntax to handle Insert, Update, and Delete without scripts, loading time was cut from over 30 minutes to 13 seconds.
To reduce costs, Kurly applied billing strategies tailored to each project's characteristics. Projects with large scan volumes reserve Reserved Slots, while projects with small scan volumes have their daily scan volume capped—significantly lowering costs compared to the previous fixed-cost model. Partitioning was also made mandatory to optimize query response times.
This adoption marks the first case in the e-Commerce industry, taking a total of 6 months from initial design to data migration. Through this effort, Kurly's Data Platform Team secured flexibility for storing large-scale data such as operation logs.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.