Impactful scheduling for GPU clusters
Allen Institute (AI2) ● Covered by 2 sources
Ai2 replaced a priority-based GPU cluster scheduler with a budgeted, hierarchical fair-share system plus time-slicing contracts to improve how it picks high-impact training jobs. The changes reduced repair work that required human involvement by 74%. As a result, GPU time is allocated via budgets (so gaming and squatting cost budget), and workloads can be preempted and requeued after minimum runtime to keep occupancy high and repairs automated.
Why it matters
We explain how Ai2’s new GPU scheduler uses time budgets, fair-share allocation, and time-slicing to prioritize high-impact research, shorten queue waits, and keep GPUs busy.