AI Briefing
KO

The Operations Layer Netflix Built Behind Large-Scale Live Streaming

·2026.04.18 00:01

Key point

Netflix has built a BOC- and TOC-centered operations system to keep pace with live scaling.

1 / 2

Details

Netflix's live operations, as recently as early 2023, was closer to ad-hoc firefighting, with engineers watching dashboards directly and coordinating over Slack. There was no dedicated operations team or formal command center, and the first live events were handled by relying on a makeshift control room and an external broadcast facility.

As the scale of live events grew rapidly afterward, the operating approach diversified in stages. The Streaming Operations Engineering(SOE) team took on live pipeline setup and in-broadcast support, easing the burden on core developers, and with the addition of a Broadcast Operations Engineer(BOE) dedicated to facility and hardware issues, a dedicated operations system extended even to physical infrastructure.

Control room operations also evolved. In the early days, a first/second captain approach had two BCOs handle a single event together, but this proved inefficient for handling around 10 concurrent events a day. As a result, the Transmission Operations Center(TOC) model was introduced, splitting operations into three roles.

  • Transmission Control Operator(TCO): manages incoming signals such as fiber, SRT, and satellite, handling up to 5 events simultaneously
  • Streaming Control Operator(SCO): manages the streaming pipeline and syndication outbound feed, handling up to 5 events simultaneously
  • Broadcast Control Operator(BCO): dedicated 1:1 to each event, covering quality, A/V sync, backup feed switching, closed captions, and SCTE messages

For the most important games or major events, the Big Bet Model is applied as an exception, assigning the entire BOC to a single event. This design is meant to secure both the efficiency of large-scale everyday operations and the stability needed for core live events where failure is not an option.

Operational reliability is tightly locked in starting from the field itself. Netflix requires 3 fully separate transmission paths for major feeds, and repeatedly runs FACS/FAX testing that includes dual power for transmission equipment, UPS, surge conditioning, facility inspections, and verification of A/V sync, latency, and captions. Ultimately, the core point of this piece is that the success or failure of large-scale live events depends not just on code, but on the human infrastructure made up of people, procedures, and facilities.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.