AI21 Labs has closed the internal chat channel where its teams used to bargain for graphics chips and now lets a waiting line decide who runs next. In a company blog post dated October 4, written by Asaf Ben-Tovim, the model developer describes the change on one shared Google Cloud cluster of roughly 10,000 GPUs (the chips that train and serve AI models). Any company with a big shared cluster has lived some version of the old arrangement.
The figure getting attention is 83 percent, and it needs careful handling. AI21 says it comes from a Google Cloud case study about the company, which makes it vendor marketing for the platform AI21 runs on, not a measurement by anyone outside the two companies. What it measures is queue wait: how much sooner the jobs ranked most urgent get going after they are submitted. It does not describe faster training, shorter runs, or a lower bill.
The post’s own table offers a possible match. It lists the average time AI21’s largest “hero” training jobs spent starved of chips falling from 72 hours to 12, which works out to about 83 percent. The post does not say the two numbers are the same measurement, and it gives no time window for either.
Before the change, engineers posted in a channel called #gpu-resources whenever the chip type they needed was taken, then hoped someone would free capacity. That worked while the cluster had spare room. Once utilization sat near 100 percent, which is what an expensive reserved fleet is supposed to do, every request turned into an argument over whose job to stop. The post lists wasted engineering time, unfair queueing, idle chips in some corners while others went hungry, and reliance on ad hoc “magic commands.”
The replacement is Kueue, open-source software that queues jobs on Kubernetes clusters, the standard system for running containerized workloads. Teams submit a job with a few labels covering priority, cost, and whether it can be interrupted. Rules, not volume, then decide the order and each team’s share. A team that has used little recently moves ahead of one that has been hogging the machines, and low-priority work can borrow idle chips and hand them back on demand.
That fairness rule has a short history. AI21 says the first design could not balance teams by chip count without also letting them preempt one another, so it raised the gap with the Kueue maintainers at Google, who shipped a feature called Admission Fair Sharing in response. The post presents this as a loop in which a heavy user’s production needs shaped an open-source roadmap.
The second problem was physical. Each machine holds eight GPUs, so scattered leftovers of 1, 1, 4, and 2 across four machines add up to eight free chips yet cannot host a job that needs all eight on one machine. A Kueue feature that knows the cluster layout now refuses jobs that cannot fit. AI21 reports fragmentation dropping from 15 percent to 8 percent.
The remaining claims are AI21’s own: manual interventions falling from 20 a week to zero, partially allocated “zombie” jobs eliminated, and the chat channel archived. All of it covers a single cluster, and the post does not say how the before numbers were measured or over what period. No independent benchmark accompanies it.
For anyone running shared GPUs near capacity, the test is cheap to copy. Count your weekly manual interventions and the hours your biggest job waits, then see whether a queue can move them anywhere near AI21’s 20 to zero and 72 to 12.
AI21 Labs, on its company blog, October 4, 2026.