Ai2 Replaces GPU Priority Labels With Time Budgets to Curb Resource Hoarding
Ai2 says automated workload draining also cut repairs needing human intervention by 74%.
Loading page…
Ai2 says automated workload draining also cut repairs needing human intervention by 74%.
Listen to this story
In an October 9 technical post, Ai2 detailed a GPU scheduler that allocates computing time through manager-set budgets instead of fixed priority labels. Teams can recover unused allocation over a rolling seven-day window, while jobs receive a declared minimum period to make progress—up to eight hours—before they may be interrupted and requeued. The rules aim to make idle capacity usable and maintenance less dependent on negotiation, while moving decisions about scarce compute into explicit research-budget choices.
Ai2 manages thousands of H100, B200 and B300 GPUs for about 150 researchers; submitted workloads request two to three times available capacity.
Jobs running on spare capacity without a budget charge can be interrupted by any budget-backed request.
Automated draining of workloads from unhealthy hosts cut repairs requiring human involvement by 74%, according to Ai2.
Ai2 has replaced its priority-based GPU scheduler with a system that gives research teams budgets of computing time rather than fixed claims on hardware. In an October 9 technical post, the institute describes rules designed to discourage resource hoarding, share capacity across projects and let running jobs be interrupted after a protected period.
The pressure is substantial. Ai2 manages thousands of NVIDIA H100, B200 and B300 GPUs, grouped into clusters of 88 to 1,024 chips. They serve about 150 internal researchers. Submitted workloads request two to three times more GPUs than are available at any given moment, the institute says.
The old rules rewarded holding on. Researchers parked idle jobs on GPUs so they could connect later, avoiding long waits for debugging work. Priority labels also lost their meaning: eventually, every scheduled workload used HIGH priority. Because users could opt out of interruption, engineers spent most of their ticket-response time negotiating shutdowns on machines needing maintenance.
Giving important projects exclusive sets of GPUs did not solve the problem. Research demand varies, so one team's reserved hardware could sit idle while another waited. Ai2 instead lets managers divide GPU time across programs, projects and researchers. Leadership decides the shares based on expected research impact; the scheduler applies those budgets as jobs arrive.
Its fair-share scheduler favors groups that have used less than their allocation over those that have used more. The default accounting window is seven days. That gives teams a way to recover time after a lull rather than lose it under a fixed limit on simultaneous GPU use. Ai2 says the scheduling algorithm itself is not new; the change is tying its weights to manager-set research budgets.
The distinction also changes the cost of hoarding. A protected idle job now spends its owner's budget on nothing. Spare capacity remains usable without a budget charge, but those jobs can be interrupted by any budget-backed request. Ai2 quotes Chris Clark saying the system feels like an extra 30% compute for his team because bursty workloads can reclaim unused allocations. That is a user's impression, not a measured increase in hardware capacity.
Budgets alone cannot balance access if a training job holds its GPUs for weeks. Ai2's scheduling contract requires each workload to declare the shortest runtime it needs to make meaningful progress. Budget-backed work is protected during that period. Afterward, it may keep running if its allocation still gives it priority, or be interrupted and automatically queued again if it is resumable.
Ai2 capped the declared minimum runtime at eight hours. Users can also choose zero, making a job free of budget charges but interruptible from the outset. The same rules let unhealthy machines shed workloads as protected periods expire, allowing repairs to proceed automatically rather than wait for negotiated shutdowns.
Ai2 acknowledges that reallocating scarce compute creates losers as well as winners. Before rollout, it built a simulator to test queue waits, interruptions and the distribution of GPU time using historical submissions and constructed scenarios. The tool let engineers vary settings such as the accounting window and minimum-runtime ceiling, simulating many days in seconds.
Budget decisions remain a human responsibility. Ai2 says researchers need frequent opportunities to argue for more time, with decisions made by managers closest to the relevant tradeoffs. The replacement moves those arguments out of individual scheduling disputes and into an explicit allocation process; it does not remove the need to choose which research deserves scarce capacity.
Loading discussion...
Join the conversation
Explain when sharing capacity should outweigh keeping a job running.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.