Same Cluster, 33 Points More Utilization: What Changed Was the Order
Dharma AI, writing on the Hugging Face blog, described a constraint-aware GPU allocator benchmarked against a FIFO scheduler across seven scenarios. On identical hardware running identical workloads, GPU utilization rose by as much as 33 percentage points, and priority-weighted output rose in every scenario, by as much as 105%.
The allocator treats real-time inference demand as a curve rather than a fixed daily reservation, allocating against demand at each timestep with batch-like work occupying troughs, bounded by a cap on how many GPUs a real-time job may swap between consecutive timesteps. Batch-like jobs are placed by priority across the whole scheduling horizon rather than in arrival order. Training, batch inference and quantization require contiguous GPU blocks held without interruption; real-time inference is elastic.
Five constraints define a legal allocation: a GPU serves at most one job per timestep; jobs respect demand ranges and inherit running work; batch-like jobs occupy contiguous blocks sized to a power of two; real-time jobs have a swap cap between timesteps; and started jobs cannot be interrupted. The objective weights real-time shortfall penalties at 5 to 10 times the batch allocation reward.
Across five contended scenarios, utilization moved from a 52–85% band to a 72–88% band, and priority-weighted value rose between 24.6% and 105.1%, averaging 52%. In a training-heavy 8-GPU case, utilization went from 53.6% to 87.0% and value rose 105%. In a scale test with 30 jobs across 64 GPUs, FIFO and the allocator both produced 44.9% utilization and completed 27 of 30 jobs, while the allocator delivered 15.9% more priority-weighted value. With all jobs set to identical priority, utilization moved from 76.8% to 87.5% and value rose 23.1%.
The system exposes a fast mode running the allocator alone on the hot path, reported at 1 to 2 milliseconds on the five contended scenarios and 15 milliseconds at 64 GPUs with 30 jobs, and a full mode using that grid as a starting point for the formal model. The scheduler optimizes a 24-hour horizon but commits only the current timestep, re-running every 30 to 60 minutes.
Based on reporting from the original publisher. Visit the source for full context and later updates.