Spotify7 min readpostmortemadvanced
AI Changed How Spotify Builds. What We Learned (and Fixed) About Quality at Higher Velocity
Summary
Spotify’s AI‑assisted development doubled change volume, exposing gaps in alerting, capacity planning, fleet‑update safety, and mobile quality signals. The team added end‑to‑end monitoring, priority‑based tiering, stronger rollback/observability, and expanded edge capacity. Data shows AI‑generated code isn’t a direct incident cause, but verification pipelines must scale with velocity.
- Add end‑to‑end failure monitoring for content ingestion to catch silent processing errors before creators notice.
- Implement workload tiering and priority‑based scheduling so critical uploads win capacity during spikes.
- Strengthen fleet‑management safeguards: expand rollback capacity, schedule automated changes during owning teams’ hours, and add more safety checks.
- Plan for compute scarcity (CPU/GPU) by reserving extra edge capacity and enabling manual traffic shifting during regional failovers.
At Spotify’s scale (3 000 services, 11 M requests/sec, 100 M concurrent clients), even minor verification gaps can affect millions. The post‑mortem shows concrete, data‑driven adjustments that keep quality high while leveraging AI‑driven velocity—a pattern other large tech orgs will need to emulate.
8/10


