AI21 Replaces Manual GPU Negotiations With Automated Queuing Across 10,000 GPUs
Kueue replaced requests in team chat, but AI21 still needed changes to share capacity fairly and ensure that free GPUs could actually fit incoming jobs.
Loading page…
Kueue replaced requests in team chat, but AI21 still needed changes to share capacity fairly and ensure that free GPUs could actually fit incoming jobs.
Listen to this story
AI21 Labs replaced its #gpu-resources channel’s capacity bargaining with Kueue-based scheduling on a shared Google Kubernetes Engine fleet of about 10,000 GPUs. The company attributes an 83% reduction in time-to-start for high-priority workloads to a Google Cloud case study. Queue rules combine team allocations, interruptible borrowing, job priority and multi-machine fit; AI21 later added fairness and topology-aware scheduling to share capacity without leaving unusable GPU fragments.
AI21 published its technical breakdown on October 4, after evaluating Apache YuniKorn and Volcano before choosing Kueue.
Kueue waits until resources are available for all coordinated container groups, preventing jobs from occupying GPUs during partial starts.
Admission Fair Sharing gives earlier admission to teams with lower historical usage, addressing a limitation in the initial allocation design.
High-priority AI workloads at AI21 Labs started sooner: the company cites an 83% reduction in time-to-start, attributing that figure to a Google Cloud case study. In a technical breakdown published October 4, AI21 explains the operational change behind that result—replacing negotiations in team chat with automated job queuing on a shared Google Kubernetes Engine cluster of roughly 10,000 GPUs.
Previously, developers asked for capacity in an internal channel called #gpu-resources and waited for someone to release it. AI21 says that approach broke down as utilization stayed near 100%. Requests became negotiations over which running work should give way, consuming engineering time while some workloads waited and pockets of capacity sat idle.
The company deliberately kept teams in one shared cluster rather than giving each its own. Its reasoning: separate Kubernetes clusters are easy to create, but sharing spare capacity between them is harder. Pooling the GPUs supported high utilization—and made allocation inside that pool more important.
AI21 chose Kueue, an open-source job-queuing system for Kubernetes, after evaluating Apache YuniKorn and Volcano. It favored Kueue’s simplicity and fit with its environment. Kueue decides when a job can start and when running work should stop to release resources.
The first design divided capacity into guaranteed allocations, interruptible work that could borrow idle resources, and on-demand machines when existing capacity could not safely be freed. Developers supplied labels for spending, interruption tolerance and priority; those labels, plus whether a job needed one machine or several, determined its queue.
Kueue also addressed a costly partial-start problem. Jobs requiring several coordinated groups of containers could begin only when resources were available for all of them, rather than occupying GPUs while waiting for the remaining pieces.
The initial implementation could not fairly share multi-machine capacity by GPU usage while protecting jobs from other teams’ interruptions except for critical-priority work. AI21 says it raised that conflict with Google’s Kueue team, helping drive the rollout of Admission Fair Sharing. That feature gives earlier admission to teams with lower historical usage.
Physical placement required a different fix. AI21 enabled Topology Aware Scheduling to account for where GPUs were available, not just their total count. The configuration could reject jobs that would not fit and pack new work onto the most-used machines first, reducing scattered spare capacity. It also avoided interruption decisions that would not actually free usable room.
Loading discussion...
Join the conversation
Explain whether past usage should influence who gets scarce resources next.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.