AI Briefing
KO

How Ccimi combined Amazon IVS with in-house infrastructure to achieve 4K, sub-4-second low-latency live streaming

·2026.05.19 17:50

Key point

Ccimi split its architecture between Amazon IVS and in-house infrastructure to achieve 4K low-latency live streaming under 4 seconds.

1 / 2

Details

To achieve both 4K, sub-4-second low latency and 1080p operational stability at the same time, Ccimi chose a hybrid structure that delegates standard areas to Amazon IVS while building only the differentiating 4K area in-house. Since the service had to be operated by a small team, general live streaming functionality was handed off to managed services, while only the core experience was directly controlled.

  • 1080p track: RTMP ingestion, transcoding, HLS delivery, automatic recording, replay, and clips were delegated to the Amazon IVS Standard channel
  • 4K track: An in-house RTMP transcoding server handles GPU transcoding and HLS serving, and the 4K stream is re-pushed to IVS to leave a recording asset

Permission verification and lifecycle events were unified through Amazon EventBridge → Amazon SQS. After a broadcast ends, IVS's S3 auto-recording is used as the source for replay and clips, and clips made during a live broadcast were stitched together by aligning the live manifest and VOD manifest with EXT-X-PROGRAM-DATE-TIME (PDT).

In a load test with 10,000 concurrent viewers before launch, about 350Gbps of traffic was needed at 4K 35Mbps, and the team also monitored p50/p95/p99 viewing latency, buffer stalled events, cache hit rate, and NAT Gateway throughput. When buffer stalled occurred for some viewers, the cause was narrowed down to CloudFront L2 fan-out, and cache keys were distributed using URL sharding, which splits requests for the same segment across multiple URLs.

For caching, a short TTL with stale-while-revalidate was applied to m3u8, while immutable long-term caching was applied to .m4s segments. On top of this, Origin Shield, startup jitter, and separation of the Accept-Encoding header were added to mitigate burst traffic, and latency and cache hit rate remained stable even through the 1,000 → 5,000 → 10,000 scaling stages.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.