AI21 Cuts High-Priority GPU Time-to-Start by 83% Using Kueue on Google Cloud
Key point
AI21 eliminated 20 weekly manual interventions and reduced hero job starvation from 72 to 12 hours by implementing Kueue on a 10,000-GPU cluster.
Details
AI21 replaced manual GPU resource negotiation with Kueue, a batch scheduler for Kubernetes, to manage a single shared GKE cluster containing approximately 10,000 GPUs. Previously, teams relied on an internal messenger channel to coordinate access, a process that resulted in zero-sum games, idle capacity, and significant engineering overhead. By adopting Kueue, AI21 automated queuing, prioritization, and cleanup, aligning resource allocation with business priorities and ensuring fair sharing across teams.
Addressing Native Kubernetes Limitations
Native Kubernetes primitives like PriorityClass and ResourceQuota failed to meet AI21's requirements for fairness and technical integrity. PriorityClass operates at the pod level, ignoring gang scheduling needs, while ResourceQuota denies admission rather than queuing, preventing idle capacity lending. Indexed Jobs also lacked all-or-nothing semantics, risking GPU waste when only some pods in a distributed training run were scheduled. Kueue was selected over alternatives like Volcano and Apache YuniKorn for its simplicity and seamless integration.
Implementation and Evolution
The rollout began with v0.10, which introduced developer labels for spend, preemptibility, and priority. However, this version suffered from a fairness gap due to the absence of Admission Fair Sharing (AFS) and a fragmentation problem where Kueue, unaware of cluster topology, could not place multi-node workloads efficiently. AI21 collaborated with the Kueue team to introduce AFS and enabled Topology Aware Scheduling (TAS) to resolve these issues.
In v0.15, AI21 redesigned the scheduling logic to include priority-aware preemptible capacity and multi-node best-effort lanes. The system now uses historical chip-hours for queue ordering and the LeastFreeCapacity algorithm to minimize fragmentation. Despite these internal changes, the developer interface remained unchanged, preserving the same three labels for routing decisions.
Operational Impact
The transition to automated scheduling yielded significant operational improvements:
- Manual interventions dropped from 20 per week to zero.
- Fragmentation decreased from 15% to 8%.
- Hero job starvation time was reduced from 72 hours to 12 hours.
- Zombie and partially allocated jobs were completely eliminated.
- The #gpu-resources channel was archived, ending the need for manual mediation of GPU disputes.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.