proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameJeffrey Dean, Sanjay Ghemawat200439 min readpaperadvanced

MapReduce: Simplified Data Processing on Large Clusters

Summary

Dean and Ghemawat present MapReduce, a simple API that lets users write map and reduce functions while the system handles parallel execution, data shuffling, and fault recovery. Their implementation processes terabytes on thousands of commodity machines, proving the model’s scalability and usability.

  • MapReduce abstracts parallelism: user‑defined map/reduce functions are automatically distributed across a cluster.
  • Fault tolerance is achieved by treating worker failures as lost tasks and re‑executing them on other nodes.
  • The shuffle phase partitions intermediate keys (hash(key) mod R) and writes them to local disks before the reduce stage.
  • Typical split size is 16‑64 MiB; the master schedules M map and R reduce tasks, enabling thousands of machines to process multi‑terabyte jobs.

Anyone building or operating large‑scale data pipelines should understand MapReduce’s model and design choices, as they underpin Hadoop, Spark, and many modern processing systems.

9/10

Related reading

  1. Bigtable: A Distributed Storage System for Structured Data

    Bigtable is a distributed storage system for structured data, designed to scale to petabytes across thousands of commodity servers. It provides a sparse, distributed, persistent multidimensional sorted map indexed by row, column, and timestamp, used by many Google products.

    Hall of Famegoogle.com48 minpaper
  2. Database Branching: A Developer's Guide to Git-Style Workflows

    Database branching uses copy‑on‑write to give developers, CI jobs, and AI agents isolated database snapshots without full copies. Branches share unchanged data, store only deltas, and are disposable, enabling production‑like testing, per‑PR isolation, safe experimentation, and rapid cleanup. Safe operation requires protecting parent branches, using mock data, TTLs, and treating migrations as the…

    Databricksdatabricks.com9 min
  3. A practical guide to cost optimization with Lakebase Postgres

    Lakebase’s separated storage‑compute architecture lets you cut database costs by syncing only the active data, picking the right sync mode, and right‑sizing compute so the hot working set fits in cache. Follow the blog’s concrete steps—materialized‑view subsets, snapshot vs triggered vs continuous sync, autoscale bounds, and connection‑pooling—to keep spend predictable without sacrificing perform…

    Databricksdatabricks.com12 min
  4. Migrating to Java 17: The Hows, Whys, and Whens

    A product‑focused overview of CloudBees CD/RO’s new integration with Argo Rollouts, describing the supported blue‑green and canary strategies, service‑mesh compatibility, and UI/analytics features. No deep technical walkthrough or performance data.

    Codeshipcloudbees.com6 min
  5. 1 points

    Saving another 100TB of RAM with math (and Rust)

    Cloudflare reduced the memory footprint of its Pingora Backend Router by re‑examining the consistent‑hashing implementation in the pingora‑ketama library. By increasing the number of virtual hash points per server from the default 1 to the standard 160 (and applying weighted hashing based on disk capacity), they cut the per‑node overhead enough to reclaim >100 TB of RAM across the fleet. The post…

    Hacker News front pagecloudflare.com13 minHN478120lobste.rs33