From manual negotiation to automated scheduling: How AI21 manages its GPU fleet with Kueue

Источник: AI21

From manual negotiation to automated scheduling: How AI21 manages its GPU fleet with Kueue

Source: AI21

An independent verifier takes agentic search past published SOTA. Our trained 8B does it for almost $0 per question.

•Updated: October 4, 2026

Google Cloud recently published our story on AI Hypercomputer, including an 83% reduction in time-to-start for high-priority workloads. Below is the full technical breakdown of how we got there with Kueue on GKE.

In brief

Sharing a large GPU fleet across many teams is a recurring challenge for anyone running AI training workloads on Kubernetes. In this blog, we share a case study from one of our shared GKE clusters, around 10,000 GPUs used across multiple teams, on how we moved from manual coordination over internal messaging channels to Kueue, an open-source, Kubernetes-native job queueing system that admits, prioritizes, and preempts workloads automatically. What began as adopting Kueue turned into a close collaboration with the team behind it at Google, with our real-world production needs feeding directly into Kueue’s roadmap. We’ll walk through what we tried, the results – a dramatic drop in metrics like manual intervention rate and the time a critical job spends “starved” – and the lessons from running it in production that fed back into Kueue itself, including new features such as Admission Fair Sharing (AFS), and is now available to all users.

A sharing problem

It’s no secret that GPUs are a scarce resource. Anyone working on model training – which requires significant GPU usage over a long period of time – is familiar with the problem up close. Their finite nature means managing GPU distribution across teams of researchers, MLOps, engineers, and alignment becomes particularly acute during training, raising challenging questions like:

  • How do you fairly share a finite, expensive resource across multiple teams?
  • Who gets prioritized when demand exceeds capacity?
  • Should you interrupt a running workload to make room for a higher-priority one?
  • How do you reconcile teams with different priorities and different definitions of “urgent”?

The “before”: Negotiating resources in team channels

At the beginning, we were navigating these questions in an internal messaging channel, called #gpu-resources. When someone needed a GPU type that was fully utilized, they asked in the channel and hoped someone would free capacity.

This kind of channel-based negotiation worked while the cluster had headroom. Once utilization stayed pinned near 100% – which is the goal for an expensive reserved GPU fleet – every incoming request became a zero-sum eviction conversation, and the channel hit a wall.

Left to manual negotiation, the cracks in this system quickly appeared:

  • Real engineering time wasted on triage and negotiation
  • Unfair queueing
  • Idle capacity in some pockets while others starved
  • Reliance on ad hoc “magic commands” to free up sufficient workload capacity

We needed to build something better. We started by devising the design principles that should inform an ideal GPU management system.

The four pillars of an ideal GPU management system

  • Fairness: Scheduling decisions are made systematically and in accordance with account usage, rather than operating on a first-come-first-served or loudest-voice-wins basis.
  • Order: Workload queuing reflects business priorities, not just timestamp.
  • Efficiency: Developers are relieved from manually deleting their idle or non-critical workloads.
  • Technical integrity: Multi-pod (gang) jobs don’t start until all resources are ready, preventing partial deployments.

From design principles to implementation: Building a GPU management system

Our building blocks

For this blog, we’ll zoom in on one of our shared GKE clusters on Google Cloud (GCP) – the one where we rolled this system out. It pools on the order of 10,000 GPUs across multiple teams.

We run it as a single shared pool rather than a cluster per team on purpose: Kubernetes makes it easy to spin up many clusters, but hard to share capacity between them. Keeping everything in one pool is what keeps utilization high – and it’s exactly what puts the pressure on how we schedule inside it.

Here’s the setup on that cluster:

We have three types of machine capacity – reserved, spot, and Dynamic Workload Scheduler’s flex-start mode– that we need to run three different types of training R&D workloads:

  • Debug pods: Developers may want to run debugging remotely on a real production environment rather than locally on their laptops
  • Training jobs: Batch workloads which can scale from 1→100s of workloads within a second. These can appear as single-node jobs or as multi-node jobs, which require multiple machines to run in parallel. We use Indexed Jobs to manage parallelism.
  • Inference models: Used by our Reinforcement Learning training jobs and managed as Kubernetes deployments, scaled up or down using HPA.

Each of these workloads can be further classified by their ability to be preempted:

  • Non-preemptible: Interrupting them mid-run is an automatic red light, as it is expensive, wasteful, and potentially damaging to time-critical workloads.
  • Preemptible: Interrupting them mid-run has minor to negligible consequences.

Searching for a fix: Evaluating vanilla K8s

Before evaluating third-party tools, we checked what Kubernetes natively ships with. Here’s what we saw:

We needed to turn to other batch scheduler options to find one that would meet all four of our criteria. As GKE doesn’t ship with a batch scheduler, we evaluated three: Apache YuniKorn, Volcano and Kueue. At that time, Volcano had more features, but its integration with our environment wasn’t seamless, and YuniKorn didn’t cover all our use cases. In terms of simplicity and integration, Kueue won.

Working with Kueue

Kueue is a set of APIs and controller for job queueing. It’s a job-level manager that decides when a job should be admitted to start (i.e. pods can be created) and when it should stop (i.e. active pods should be deleted).

Step 1: Validating system criteria

Mapping our four pillars to Kueue’s features, we saw it could meet all of our criteria:

  • Fairness: Kueue offers two complementary mechanisms: Preemption Fair Sharing (preempts workloads to balance cluster usage) and Admission Fair Sharing (AFS) (admits workloads first from teams that historically used less).
  • Order: Workload Priority Classes close the gap we identified in native Kubernetes PriorityClasses, controlling both queue order and preemption order. Admission Fair Sharing adds another lever on top of timestamp-based ordering.
  • Efficiency: Kueue’s core resource, the ClusterQueue, gives admins fine-grained control over preemption behavior and integrates Fair Sharing + Workload Priority Classes. Kueue handles workload cleanup automatically.
  • Technical integrity: Kueue’s all-or-nothing semantics integrate cleanly with our Indexed Jobs: multi-pod gangs are only admitted when every pod can run.

Step 2: Defining implementation requirements

Kueue supported the overarching behaviors we needed for our internal GPU management system. Yet, to guide our implementation, we would need to map some more requirements. We sat down with our primary stakeholders – the Algorithm Developer Leads who manage cluster usage – to pin down the exact behavior they needed from the system:

Step 3: Rolling out our first design (v0.10)

From the admin perspective, we built three layers of capacity into our first version:

  • Guaranteed capacity: Each team has its own quota for single-node workloads, a shared pool for multi-node workloads, and another shared pool for single-node workloads.
  • Preemptible capacity: A single shared pool with fair-sharing between tenants, able to borrow from any guaranteed capacity that was idle. This is what keeps utilization high and helps distribute capacity across teams.
  • On-demand capacity: For when the system was full and no workload could be safely preempted: developers could provision new machines dynamically. Also useful for testing GPU types not in the reserved pool.

From the developer’s perspective, we wanted routing to a queue to be as straightforward as possible. Choose labels to describe workload spend (low/high), ability to be preempted (false/true), and priority (low/medium/high/critical); those answers, along with whether the workload is single-node or multi-node, would determine which queue it ultimately lands in.

Where v0.10 failed: Fairness and fragmentation

The fairness gap

At the time we implemented this solution, Admission Fair Sharing (AFS) didn’t yet exist, and we were limited to using preemption Fair Sharing in our system. This blocked us from achieving these two multi-node requirements at the same time:

  • Fair scheduling between tenants based on chips
  • No preemption in this quota between different teams unless someone explicitly specifies “critical priority” for their indexed job

In keeping with the open source ethos of contribution, we raised this challenge directly with the Kueue team; in response, they quickly rolled out AFS, which enables two critical behaviors:

  • Preempts workloads from a team that starves other teams
  • Changes the order of the workloads in the queue to prioritize low usage teams

For us, this represents the best case scenario feedback loop that fuels any strong open source project: uncovering usage requirements through real-world usage and feeding that knowledge back into the product.

It also meant we now had a solution to the fairness gap.

The fragmentation problem

We also ran into a second issue: Sometimes pods got admitted while no single machine could actually fulfill the pod’s requirements. Our nominal quotas sum to the total number of GPUs we have, so at first this looks like a bug.

Here’s the case that exposed it: Each machine has 8 GPUs, and we submit a workload that needs all 8 on one node. If GPU capacity is scattered as 1+1+4+2 across four nodes, the quota math says 8 GPUs are free — but an 8-GPU job has nowhere to land. That is the kind of fragmentation we hit.

While GKE does its best to bin-pack, Kueue isn’t the scheduler in our setup, and it isn’t aware of our cluster topology (8 GPUs per machine).

The fix we found was a Kueue feature we hadn’t utilized in v0.10: Topology Aware Scheduling (TAS). While TAS is frequently activated around gang scheduling – e.g. keeping distributed workloads physically nearby to reduce latency and cost – our motivation was different.

We frequently had more than 8 GPUs free across the cluster but no single node with 8 free. A straightforward custom resource called Topology lets admins tell Kueue which topologies to consider when making scheduling decisions. Technically Kueue isn’t a scheduler, but when TAS is configured Kueue sets nodeName on admitted workloads. For our purposes, we now had a new scheduler.

With AFS and TAS in hand, we reconvened with our primary stakeholders, the Algorithm Developer Leads. The original backlog was now solvable:

  • AFS permitted fair multi-node sharing
  • And TAS meant Kueue could refuse to admit pods that wouldn’t fit; schedule with LeastFreeCapacity algorithm, which packs new workloads onto the most-utilized nodes first to minimize fragmentation; and stop making preemption decisions that don’t actually free capacity.

We had now solved the fairness gap and the fragmentation problem.

But as in any good meeting, we walked out of the conversation with two new requests from our stakeholders:

  • Priority-driven preemption inside the preemptible single-node pool. A priority=low single-node preemptible workload can be preempted by medium/high workloads in the same pool.
  • A new multi-node best-effort lane. Multi-node workloads can opt in with priority=low; they get preempted by medium/high multi-node workloads when capacity is tight, and they still preempt single-node preemptible workloads when needed.

Given what we had learned about AFS and TAS, as well as the two new asks, we re-designed our Kueue implementation to a v0.15 version.

Redesigning to include AFS + TAS (v0.15)

Our updated version, v0.15, preserved the v0.10 layout yet included two adaptations to support the required behaviors we had identified in our first rollout:

  • Preemptible capacity is now priority-aware. The shared single-node preemptible pool splits by workload priority — priority=low workloads run in their own lane and get preempted first when medium/high workloads need room. A new multi-node best-effort lane joins the design alongside it: multi-node workloads can opt in with priority=low, where they preempt single-node preemptible workloads when needed and themselves get preempted by medium/high multi-node workloads.
  • Admission and placement got smarter. Admission Fair Sharing (AFS) now rebalances queue order based on historical chip-hours used per team, on top of priority and timestamp ordering. Topology Aware Scheduling (TAS) refuses to admit workloads that wouldn’t fit on a single node and schedules with LeastFreeCapacity to keep the cluster packed.

Guaranteed Capacity and On-Demand Capacity remain unchanged from v0.10.

The developer-facing decision tree picks up the new lane:

From the user’s perspective, there aren’t new labels nor new label-values, so the interface didn’t change. The only thing that is changed is how priorities affect the routing of the workload – low is now separated from medium and high for multi-node workloads and preemptible single-node workloads

“The after:” The impact of implementing Kueue

GPUs are no longer negotiated via internal team channels. Today, every workload at AI21 – whether single-node debug pod, multi-node training run, or inference deployment – gets automatically assigned to the right queue. Kueue automatically manages the process from end to end, no manual negotiation required: from queuing and prioritization, to cleanups and fairness of GPU usage across teams.

Today, the #gpu-resources channel is archived. Group leads aren’t refereeing GPU fights, and researchers spend their time on experiments instead of asking for capacity. Good plumbing isn’t just GPU hours – it’s our researchers being able to focus on what they love and excel at: the science. Here’s what we gained back by the numbers once we switched to Kueue:

Operational wins

  • Manual interventions per week: 20 → 0
  • “Zombie” / partially-allocated jobs eliminated
  • Fragmentation percentage: 15% → 8%
  • Time wasted in team messaging channels for GPU negotiation per week: hours → none
  • Average time a hero job gets starved: 72 hours → 12 hours

Takeaways & conclusion

For anyone tackling the challenge of sharing GPU resources within a cluster, here are the core lessons we learned as we transitioned from a manual negotiation system to automatic management with Kueue:

  • Schedule at the workload level, not the pod level. Native Kubernetes preemption will happily kill one pod of a four-pod training run and leave the other three burning GPUs. Distributed training demands gang/all-or-nothing admission.
  • A queue isn’t a scheduler until it understands your topology. The quota math said 8 GPUs were free; fragmentation (1+1+4+2 across four nodes) meant an 8-GPU job had nowhere to land. Topology Aware Scheduling is what let Kueue refuse jobs that couldn’t fit and pin placement onto the most-utilized nodes first, turning a queue into something that actually schedules.
  • Keep the fleet pinned near 100% on purpose. For a reserved GPU fleet, full utilization is the target, not an alarm. What makes that safe is a preemptible/opportunistic lane that borrows idle capacity and gives it back on reclaim, so “full” doesn’t translate to “blocked.”

As any company training models can tell you, sharing GPUs across teams isn’t an easy problem to solve. When usage is negotiated asynchronously over internal team channels or divided decisively into per-team clusters, it leaves much to be desired in terms of time and compute efficiency. Both methods risk idle GPUs, inconsistencies with business priorities, and interrupted jobs.

Only once we shifted to using Kueue to handle admission, priority, and preemption for our team’s single shared cluster, capacity could be allocated automatically: fair across teams, ordered by business priority, and reclaimed without manual intervention.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.