Google Cloud BlogBarak Peleg5 min readintermediate
AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer
Summary
AI21 moved from ad‑hoc Slack requests to an automated scheduling stack using GKE and the Kueue scheduler on Google Cloud AI Hypercomputer. The change cut high‑priority job start‑up latency from 72 h to 12 h, eliminated manual interventions, and reduced GPU fragmentation, all without increasing cost.
- Switching to Kueue on GKE reduced high‑priority job wait time from 72 h to 12 h (≈83% faster).
- Manual scheduling interventions dropped from ~20 per week to zero, freeing engineering time.
- Admission Fair Sharing and Topology‑Aware Scheduling cut GPU fragmentation from 15% to 8% and removed zombie jobs.
- Integration with AI Hypercomputer allowed automatic spill‑over to Spot VMs and Dynamic Workload Scheduler without rewriting specs.
Teams running large multi‑node GPU training on shared clusters need automated, fair scheduling to avoid bottlenecks and manual coordination.
6/10





