InfoQOlimpiu Pop2 min readintermediate
From Memory-Hungry HNSW to Quantized SPANN: The Technical Evolution of Pinterest's Manas Platform
Summary
Pinterest reengineered its Manas search platform to replace memory‑heavy HNSW with scalar and product quantization, cutting index size by up to 93% while keeping recall above 70% and achieving 20‑30% cost savings. They also moved ANN storage to SSD using SPANN, gaining 3× query throughput over DiskANN with modest latency and recall impact.
- Product Quantization shrinks HNSW from 121 GB to 32 GB (‑74%) with recall ~77% and similar QPS.
- Scalar Quantization reduces HNSW to 50 GB (‑59%) while preserving >90% recall and query speed.
- IVF + SQ yields the best recall (95.7%) on a 25 GB index with high QPS.
- SPANN on SSD delivers 3× the QPS of DiskANN, 1/3 the latency, and only ~5% recall loss, saving ~40% CPU time.
Engineers building billion‑scale vector search systems need practical techniques to cut memory and cost while maintaining performance, and Pinterest's quantization and SSD‑backed SPANN provide proven solutions.
6/10


