AI Briefing
KOSign in

Airbnb Builds System to Capture and Replay Real Database Traffic for Load Testing and Upgrades

·2026.10.07 02:01

Key point

The system enabled Airbnb to migrate its MySQL fleet from 5.7 to 8.0 without major production incidents by identifying latency regressions and compatibility issues offline.

1 / 4

Details

Airbnb replaced its fragmented, client-side query logging system with a unified infrastructure built on ProxySQL to capture and replay real production database traffic. This new system addresses the limitations of synthetic benchmarks by providing a complete picture of transactions, enabling accurate load testing, capacity planning, and compatibility verification for database upgrades.

System Architecture

The solution consists of three in-house components that process traffic flowing through ProxySQL:

  • Log Mover: A sidecar that transfers ProxySQL query logs from local disk to cloud object storage.
  • Log Processor: An offline job that decodes binary logs, partitions queries by cluster, reassembles transactions in original order, and buckets them into five-minute windows. It also rewrites INSERT statements to pin last_insert_id values, ensuring consistent behavior across different database engines during replay.
  • Log Replayer: A distributed system with an API server, task scheduler, and horizontal worker fleet. It supports two modes: Replay only for load testing at configurable speeds (e.g., 2x traffic) and Replay and compare for running identical queries against two targets to detect discrepancies.

Impact on Upgrades and Capacity Planning

The system proved critical during Airbnb’s migration from MySQL 5.7 to 8.0. Replay testing identified specific performance regressions, such as a query pattern where latency increased from 0.03 to 2.6 seconds due to changes in row sorting in MySQL 8.0.20+. It also surfaced correctness issues, such as non-deterministic row ordering when LIMIT was applied without a unique tiebreaker, which was resolved by adding explicit ordering.

For capacity planning, the team replays production traffic at higher multiples to model future growth. In one instance, replaying 80% more write traffic on a large cluster increased average commit latency from 6 ms to 34 ms, revealing the cluster's ceiling before real-world traffic hit it. This allows teams to right-size clusters and optimize costs ahead of peak seasons.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.