proomt

Search

Search posts, papers, and topics

All posts

InfoQSteef-Jan Wiggers4 min readintermediate

GKE Pod Snapshots Cut Model Load Times, and Move the Work to Snapshot Lifecycle Management

Summary

GKE Pod snapshots reduce AI model load times by up to 89%, enabling 70B models to load in 37s and 8B models in 15s. This feature checkpoints and restores the entire running state, including memory, but requires gVisor and shifts complexity to snapshot lifecycle management.

  • GKE Pod snapshots save and restore the full running state, including CPU/GPU memory, bypassing model initialization.
  • It requires GKE Sandbox (gVisor) and is configured via PodSnapshotStorageConfig and PodSnapshotPolicy CRDs.
  • Startup latency for large AI models can be cut significantly, e.g., 70B model from minutes to 37 seconds.
  • Snapshot invalidation is a key challenge, as changes in Pod spec, gVisor, or GPU drivers can render them unusable.

Engineers running large AI inference models or other stateful workloads on GKE should care, as this feature can drastically reduce startup times and improve resource utilization.

7/10

Related reading

  1. Agent Substrate brings high-density, scalable, trusted infrastructure to GKE

    Agent Substrate is an open‑source runtime for AI agents that runs on GKE. It uses Cloud Hypervisor microVMs or gVisor sandboxes to give kernel‑level isolation, a custom control‑ and data‑plane that can suspend/resume agents in <500 ms, and a “zero‑idle” model that packs >1 000 dormant agents per host (≈10× density vs. containers). GKE integration adds custom ComputeClasses, spot/on‑demand pools,…

    Google Cloud Bloggoogle.com6 min
  2. Kubernetes v1.37: Pod-Level Resource Managers graduated to Beta

    Kubernetes v1.37 adds Pod‑Level Resource Managers to beta (off by default). The feature lets Kubelet’s Topology, CPU, and Memory managers consume pod‑level `.spec.resources` to reserve exclusive NUMA‑aligned CPUs/memory for primary containers while sidecars share a pod‑isolated pool. A new PodResources gRPC API now reports `cpu_ids` and `memory` per pod. Enable via the `PodLevelResourceManagers`…

    Kuberneteskubernetes.io2 min
  3. For SeaVerse, GKE Agent Sandbox reduces infrastructure costs by 60%

    SeaVerse uses GKE Agent Sandbox (Kata Containers + Cloudhypervisor or gVisor) to run isolated AI sandboxes at scale, achieving 300 allocations / s per cluster (90% ≤ 200 ms) and cutting infrastructure spend by up to 60% via flexible VM sizing and per‑sandbox persistent storage, while gaining native Cloud observability.

    Google Cloud Bloggoogle.com5 min
  4. What’s new in AI infrastructure and orchestration in September

    This Google Cloud blog post details September updates to its AI infrastructure and orchestration, focusing on scalability for 'agentic' AI workloads. Key enhancements include new GKE features like Agent Substrate for high-density sandboxes, native scale-to-zero, and Pod snapshots, alongside storage improvements like Filestore agent volumes and new M4N/Z4D VMs.

    Google Cloud Bloggoogle.com15 min
  5. Monitor TAS and gang scheduling for AI training in Kubernetes

    Kubernetes’ default scheduler can’t satisfy AI training’s need for low‑latency GPU interconnects and simultaneous pod start‑up. The blog explains how the open‑source Kueue job queue adds topology‑aware placement (using node labels like `topology.kubernetes.io/rack`) and how the Coscheduling plugin adds a permit phase that only binds a gang of pods when the full set is ready, preventing idle GPU r…

    Datadogdatadoghq.com19 min