Impactful scheduling for GPU clusters

The AI Infrastructure team at the Allen Institute for AI (AI2) has replaced its priority-based cluster scheduler with a budget-driven system combining GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The change was rolled out cluster by cluster starting at the end of July. AI2 manages thousands of NVIDIA H100, B200, and B300 GPUs across clusters ranging from 88 to 1024 GPUs, serving roughly 150 internal researchers working on LLM and VLM training, robotics reinforcement learning simulation, and post-training for scientific agentic use cases.
The problem AI2 set out to solve was structural, not a matter of tuning priorities. Demand for GPU time runs 2-3x above supply, meaning every available GPU hour has two or three competing research workloads. Under the old scheduler, teams could opt workloads out of preemption and held limits on concurrent GPUs for those protected jobs, while preemptible work could use idle capacity beyond the limit. That produced two recurring pathologies: GPU squatting, where users parked no-op workloads they could connect to later, and priority inflation, where eventually 100% of scheduled workloads were marked HIGH, starving lower tiers of time entirely. AI2 also says on-call engineers spent most of their ticket time negotiating shutdowns of non-preemptable workloads on hosts with known maintenance problems.
AI2 frames this as a tragedy-of-the-commons pattern and cites the 2011 Dominant Resource Fairness paper by Ghodsi et al., which recounts a search company that gave dedicated machines to jobs whose users could guarantee high utilization, only to find users adding infinite loops to inflate measured utilization. The institute's earlier fix, granting teams monopolies over blocks of GPUs, was too coarse: because research demand is seasonal, monopoly holders sometimes had nothing to run while other teams waited, leaving capacity idle.
The replacement allocates GPU time rather than GPUs. Managers proportionally distribute shares through a hierarchy that mirrors the research program structure, so a project might hold a 35% claim on total capacity. Every request for protected GPU time must be funded by a budget or it is not shielded from preemption. Allocation decisions sit with whoever has the most context on the tradeoffs: lead researchers within a project, principal investigators within a program, and a lead program manager or the CEO across programs.
The scheduler underneath is hierarchical fair-share, an approach AI2 notes dates back to the Hadoop Fair Scheduler in 2009 and remains in use in SLURM's Fair Tree and YARN's Fair Scheduler. What differs is the input: the tree mirrors the program structure and weights come from manager-set budgets rather than static quotas. Occupancy is tracked over a sliding lookback window, defaulting to 7 days, and workloads from under-utilized allocations are sorted above those from over-utilized ones. The system distinguishes allocated occupancy, charged to a budget and protected during a minimum runtime window, from unallocated occupancy, which is unprotected, uncharged, and preemptible by any allocated request. That second category keeps GPUs busy when allocations do not match demand and means teams never need to decline free cycles.
The time-slicing contract addresses long-running distributed training, where jobs can occupy GPUs for days or weeks. In exchange for cluster access, a workload declares its minimum runtime, the shortest occupancy needed for meaningful progress, and is protected from preemption for that period. Afterward the scheduler may rebalance and automatically requeue resumable work. Setting minimum runtime to zero marks the GPU time as unallocated. A workload lifecycle thus runs: submission with minimum runtime and resumability flag; fair-share scheduling; protected minimum runtime charged to allocations; continued running while allocations still prioritize it; possible preemption and requeue; and completion. This also lets unhealthy hosts drain workloads as they hit minimum runtime, enabling automated repair. AI2 says human-in-the-loop repairs fell 74%, a larger saving than anticipated.
Before rollout, AI2 built a simulator that takes workloads and their submission schedule, makes preemption and GPU assignment decisions, and can jump ahead to schedulable moments, producing queue wait and preemption analysis over many simulated days in seconds. Configurable knobs included lookback window length and maximum minimum runtime, which was set to 8 hours. A key hypothesis concerned debug workloads, small jobs with minimum runtimes of 15 minutes or less that test whether a job launches or crashes early; a minute or two of wait would enable a new development practice, while ten minutes would not. Historical data lacked enough such jobs, so test cases were hand-built; simulations predicted p90 debug wait times falling from about 6 hours to 5 minutes.
Real results beat predictions. Over a 30-day test period, teams received 98% of the GPU hours they were owed, counting time owed as allocation capped hourly at actual demand; 13 of 15 team allocations received 95% or more, and the worst case received 90%. Occupancy held steady at 98% before and after, with demand still 2-3x capacity. Unallocated time accounted for 18% of delivered GPU hours. Debug workload p90 queue wait fell from 2 hours to 30 seconds. On the largest H100 cluster, median queue wait dropped from 5 minutes to 24 seconds and p90 fell by about a third, from 2.8 hours to 1.8 hours. On the three original problems, AI2 reports debug jobs starting in under a minute, priority now sorting only within a team, and the 74% reduction in human-in-the-loop repairs.
The rollout was not frictionless. Because it proceeded incrementally, early behavior varied by cluster, and interfaces kept old terminology such as workload priority whose meaning had changed. Documentation alone did not clear up the confusion; live explanatory sessions with real examples did, and AI2 describes a shift from frustration and folk theories toward research groups communicating more openly about GPU needs and engaging in budgeting discussions with clearer knowledge of tradeoffs.
Why it matters: teams running shared GPU fleets face the same economics AI2 describes, and this is a concrete, published account of replacing discretionary priority with funded allocations plus short protected runtimes, including measured numbers rather than intentions. The 74% reduction in manual repair work and the sub-minute debug job latency are the kind of operational gains that matter as much as raw throughput, because they determine how quickly researchers can iterate. Inference: the generalizable ideas here are not the specific tree or lookback window but the principle that free priority and optional preemptability invite gaming, and that charging for protection while leaving everything else preemptible tends to keep utilization high without hard partition. The cost is administrative: budgets require ongoing review, and the reported learning curve suggests documentation alone will not carry a migration like this.
Based on reporting from the original publisher. Visit the source for full context and later updates.