proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameGiuseppe DeCandia et al.200765 min readpaperadvanced

Dynamo: Amazon's Highly Available Key-value Store

Summary

Dynamo is Amazon's highly available key‑value store that trades strong consistency for availability using consistent hashing, vector‑clock versioning, quorum reads/writes, and a gossip‑based membership protocol. The paper shows that an eventually‑consistent design can run at massive scale with strict latency SLAs.

  • Uses consistent hashing for partitioning and replica placement, enabling seamless scaling and node churn.
  • Employs vector clocks and client‑side conflict resolution to handle concurrent writes without locking.
  • Read/write quorum (R+W>N) provides tunable consistency guarantees per application.
  • Gossip protocol disseminates membership and failure information without a central coordinator.

Engineers building distributed storage or services that need high availability should understand Dynamo's trade‑offs and techniques, which underpin many modern NoSQL systems.

8/10

Related reading

  1. Amazon Aurora: Design Considerations for High Throughput Cloud-Native Relational Databases

    Amazon Aurora is a cloud-native relational database that decouples compute from storage, offloading redo log processing to a purpose-built, multi-tenant storage service. This architecture significantly reduces network I/O, enables fast crash recovery, and provides high durability and availability through a novel 6-way replication quorum model.

    Hall of Famestanford.edu47 minpaper
  2. WAL + S3: Lakebase storage for the era of agents

    Neon's Lakebase Postgres redefines database storage by making the WAL the source of truth, stored on S3, rather than data files. This transaction-centric approach enables instant branching, time travel, and scalable, decoupled compute and storage for Postgres.

    Neonneon.com14 minHN4
  3. A practical guide to cost optimization with Lakebase Postgres

    Lakebase’s separated storage‑compute architecture lets you cut database costs by syncing only the active data, picking the right sync mode, and right‑sizing compute so the hot working set fits in cache. Follow the blog’s concrete steps—materialized‑view subsets, snapshot vs triggered vs continuous sync, autoscale bounds, and connection‑pooling—to keep spend predictable without sacrificing perform…

    Databricksdatabricks.com12 min
  4. Accelerating the borderless Lakehouse: Announcing preview of cross-cloud caching

    Google Cloud previewed cross‑cloud caching for its Borderless Lakehouse. The feature caches sub‑file Parquet blocks in Google Cloud, encrypts them with GMEK, isolates cache per tenant/region, and validates freshness via metadata checks. In tests it can reduce cross‑cloud data transfer to <5% of the original size, lowering query latency and cost for Iceberg tables stored in other clouds. BigQuery…

    Google Cloud Bloggoogle.com3 minrelease