proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameFay Chang et al.200648 min readpaperintermediate

Bigtable: A Distributed Storage System for Structured Data

Summary

Bigtable is a distributed storage system for structured data, designed to scale to petabytes across thousands of commodity servers. It provides a sparse, distributed, persistent multidimensional sorted map indexed by row, column, and timestamp, used by many Google products.

  • The data model is a sparse, distributed, persistent multidimensional sorted map (row, column, timestamp) -> string.
  • Rows are lexicographically ordered and dynamically partitioned into "tablets," which are units of distribution and load balancing.
  • Column keys are grouped into "column families" for access control, type grouping, and storage optimization.
  • Each cell can store multiple versions, indexed by timestamp, with configurable automatic garbage collection policies.

This foundational paper introduced a key-value store architecture that influenced many subsequent NoSQL databases, making it essential for engineers designing or working with large-scale distributed data systems.

9/10

Related reading

  1. The Google File System

    The Google File System (GFS) is a scalable distributed file system designed for Google's data-intensive applications, built on inexpensive commodity hardware. It provides fault tolerance and high aggregate performance by optimizing for large files, sequential appends, and anticipating frequent component failures.

    Hall of Famegoogle.com62 minpaper
  2. Introducing Filestore agent volumes: fully managed storage for agent workspaces

    Google Cloud adds Filestore agent volumes, a fully‑managed, elastic file‑system that automatically provisions isolated POSIX workspaces for GKE‑based AI agent sandboxes. Volumes attach in milliseconds, support RWX with file‑level locking, and charge only for used capacity with automatic tiering, aiming to cut cold‑start latency and storage waste for large‑scale agent fleets.

    Google Cloud Bloggoogle.com4 min
  3. Architecture of a Database System

    The paper surveys the internal architecture of modern relational DBMSs, describing the main components—client communications, process management, query processing, storage, and transaction management—and how they interact during a query lifecycle. It highlights design decisions such as admission control, buffer management, and concurrency control that are critical for scalability and reliability.

    Hall of Fameberkeley.edu161 minpaperHN31
  4. Spanner: Google's Globally-Distributed Database

    Spanner is Google's globally-distributed, synchronously-replicated database, notable for being the first to support externally-consistent distributed transactions at global scale. Its key innovation is the TrueTime API, which exposes clock uncertainty to enable these strong consistency guarantees.

    Hall of Famegoogle.com46 minpaper
  5. MapReduce: Simplified Data Processing on Large Clusters

    Dean and Ghemawat present MapReduce, a simple API that lets users write map and reduce functions while the system handles parallel execution, data shuffling, and fault recovery. Their implementation processes terabytes on thousands of commodity machines, proving the model’s scalability and usability.

    Hall of Famegoogle.com39 minpaper