proomt

Search

Search posts, papers, and topics

All posts

Google Cloud BlogBarak Peleg5 min readintermediate

AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer

Summary

AI21 moved from ad‑hoc Slack requests to an automated scheduling stack using GKE and the Kueue scheduler on Google Cloud AI Hypercomputer. The change cut high‑priority job start‑up latency from 72 h to 12 h, eliminated manual interventions, and reduced GPU fragmentation, all without increasing cost.

  • Switching to Kueue on GKE reduced high‑priority job wait time from 72 h to 12 h (≈83% faster).
  • Manual scheduling interventions dropped from ~20 per week to zero, freeing engineering time.
  • Admission Fair Sharing and Topology‑Aware Scheduling cut GPU fragmentation from 15% to 8% and removed zombie jobs.
  • Integration with AI Hypercomputer allowed automatic spill‑over to Spot VMs and Dynamic Workload Scheduler without rewriting specs.

Teams running large multi‑node GPU training on shared clusters need automated, fair scheduling to avoid bottlenecks and manual coordination.

6/10

Related reading

  1. Monitor TAS and gang scheduling for AI training in Kubernetes

    Kubernetes’ default scheduler can’t satisfy AI training’s need for low‑latency GPU interconnects and simultaneous pod start‑up. The blog explains how the open‑source Kueue job queue adds topology‑aware placement (using node labels like `topology.kubernetes.io/rack`) and how the Coscheduling plugin adds a permit phase that only binds a gang of pods when the full set is ready, preventing idle GPU r…

    Datadogdatadoghq.com19 min
  2. From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

    NVIDIA’s DSX platform lets AI data‑centers shift workloads in response to grid signals, squeezing ~24% more token throughput (4 M→5 M tps) and ~23% better performance‑per‑watt on a fixed megawatt budget. The first production demo used Emerald AI’s Conductor to drop a 4 MW load to 3 MW in under a minute without interrupting high‑priority jobs. DSX MaxLPS reallocates headroom across HGX B200 server…

    Nvidianvidia.com5 min
  3. How to upskill enterprise AI builders by using daily micro habits

    Google Cloud Consulting proposes a four‑pillar micro‑learning framework for enterprise AI upskilling: 5‑minute browser‑based exercises, pre‑configured sandboxes, daily streaks, and delivering runnable code each session. A pilot (Advent of Agents) showed >150k participants, 859k code runs, and a 31% daily return rate, suggesting short, frictionless tasks improve engagement versus traditional bootc…

    Google Cloud Bloggoogle.com3 min
  4. What’s new in AI infrastructure and orchestration in September

    This Google Cloud blog post details September updates to its AI infrastructure and orchestration, focusing on scalability for 'agentic' AI workloads. Key enhancements include new GKE features like Agent Substrate for high-density sandboxes, native scale-to-zero, and Pod snapshots, alongside storage improvements like Filestore agent volumes and new M4N/Z4D VMs.

    Google Cloud Bloggoogle.com15 min