Ai2 Replaces Priority Scheduler with GPU Time Budgets, Cutting Debug Wait Times from 2 Hours to 30 Seconds
Key point
The new system reduced p90 debug workload wait times from 2 hours to 30 seconds and cut on-call maintenance toil by 74% across clusters of thousands of NVIDIA GPUs.
Details
Ai2’s AI Infrastructure team replaced its priority-based GPU scheduler with a system based on GPU time budgets, hierarchical fair-share, and time-slicing contracts to address severe demand-supply imbalances where GPU demand exceeded capacity by 2–3x. The change aims to convert opaque resource contention into a transparent administrative budgeting process, preventing issues like "squatting" and priority inflation.
Core Mechanisms
The new architecture shifts from GPU ownership to time allocation, structured hierarchically from programs to projects to researchers.
- Fair-Share Scheduler: Uses a sliding lookback window (default 7 days) to track occupancy. It distinguishes between Allocated occupancy (budget-deducted, protected by minimum runtime) and Unallocated occupancy (no budget deduction, preemptible, used to fill idle cycles).
- Scheduling Contracts: Workloads must declare a minimum runtime to prevent preemption during critical execution. After this period, jobs can be preempted and requeued to ensure fairness.
- Budget Enforcement: Leadership sets GPU time budgets for research projects. Requests exceeding budgets are preemptible, increasing the cost of scheduler abuse.
Measured Results
Following a 30-day test rollout starting in late July, the system delivered significant operational improvements:
- Debug Workload Latency: p90 queue wait time for debug workloads dropped from 2 hours to 30 seconds (simulation predicted 6 hours to 5 minutes).
- Cluster Occupancy: Maintained 98% occupancy before and after the change, despite demand exceeding capacity by 2–3x.
- Budget Adherence: Teams received 98% of their allocated GPU hours; 13 of 15 teams received over 95%.
- On-Call Toil: Automated workload draining and restarting reduced human-in-the-loop maintenance requirements by 74%.
- Queue Latency: Median queue wait time for the largest H100 cluster fell from 5 minutes to 24 seconds, with p90 wait times decreasing by roughly one-third.
Challenges and Roadmap
The transition introduced a steep learning curve and cultural friction, requiring live explanatory sessions to clarify new terminology. A key limitation emerged for interactive sessions (e.g., data analysis), where volatile state dependencies made preemption costly compared to the previous system's ability to hold resources for up to a week. Ai2 is addressing this by investing in CPU-only clusters for data prep and developing restorable sessions for CPU workloads. The team is also investigating capacity fragmentation, where minimum runtime protections may be increasing wait times for the largest workloads by reducing simultaneous preemption opportunities.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.