proomt

Search

Search posts, papers, and topics

All posts

PlanetScaleSimeon Griggs6 min readintermediate

Handling hot shards

Summary

This article explains how PlanetScale's Neki, built on Vitess, helps manage hot shards in sharded databases, using Slack's journey from workspace-based to channel-based sharding as a detailed case study. It outlines strategies like vertical scaling, tenant isolation, and resharding specific tables without downtime to adapt to changing data access patterns.

  • Sharding by `tenant_id` can lead to hot shards if a single tenant grows excessively, requiring adaptation.
  • Temporary fixes for hot shards include vertical scaling the affected shard or isolating the hot tenant to its own shard.
  • The long-term solution often involves resharding specific tables with a more appropriate key based on access patterns, like sharding messages by `channel_id` instead of `workspace_id`.
  • Neki/Vitess supports declarative data topologies and online resharding, allowing changes to sharding strategies without application downtime.

Engineers designing or operating large-scale sharded databases will find practical strategies and a detailed real-world example for managing uneven data distribution and evolving access patterns.

7/10

Related reading

  1. The architecture of Neki

    Neki is PlanetScale’s sharding layer for vanilla PostgreSQL that presents a single Postgres endpoint while routing queries across a fleet of Postgres instances. It does this with a set of tightly‑coupled components—Router, Sidecar, PostgresManager, Admin, Operator, and etcd‑backed Data Topology—each handling a specific piece of the scaling, failover, and query‑planning puzzle.

    PlanetScaleplanetscale.com8 min
  2. A practical guide to cost optimization with Lakebase Postgres

    Lakebase’s separated storage‑compute architecture lets you cut database costs by syncing only the active data, picking the right sync mode, and right‑sizing compute so the hot working set fits in cache. Follow the blog’s concrete steps—materialized‑view subsets, snapshot vs triggered vs continuous sync, autoscale bounds, and connection‑pooling—to keep spend predictable without sacrificing perform…

    Databricksdatabricks.com12 min
  3. Designing Neki for performance

    Neki, a PostgreSQL proxy, achieves high performance by lazily decoding query results and bind parameters. It avoids unnecessary parsing and data conversions, keeping data in raw PostgreSQL wire format until specific operations like sorting or limiting require deeper inspection.

    PlanetScaleplanetscale.com15 minHN2
  4. We're making Tailscale faster

    Tailscale is cutting memory overhead for small packets, adding a multi‑queue pipeline for routers and exit nodes, using Linux’s writev, and introducing netmap caching to speed up startup. These changes give ~5 % throughput gains now and larger gains in upcoming releases.

    Tailscaletailscale.com7 minHN242101
  5. Blocking cutovers to save replication slots

    PlanetScale’s Postgres operator blocks cutovers that would leave logical replication slots out‑of‑sync. By requiring `failover=true` on slots and turning on `hot_standby_feedback` and `sync_replication_slots`, the operator ensures slots survive a primary promotion, preventing downstream CDC consumers from losing data.

    PlanetScaleplanetscale.com5 min