Hall of FameJeffrey Dean, Sanjay Ghemawat200439 min readpaperadvanced
MapReduce: Simplified Data Processing on Large Clusters
Summary
Dean and Ghemawat present MapReduce, a simple API that lets users write map and reduce functions while the system handles parallel execution, data shuffling, and fault recovery. Their implementation processes terabytes on thousands of commodity machines, proving the model’s scalability and usability.
- MapReduce abstracts parallelism: user‑defined map/reduce functions are automatically distributed across a cluster.
- Fault tolerance is achieved by treating worker failures as lost tasks and re‑executing them on other nodes.
- The shuffle phase partitions intermediate keys (hash(key) mod R) and writes them to local disks before the reduce stage.
- Typical split size is 16‑64 MiB; the master schedules M map and R reduce tasks, enabling thousands of machines to process multi‑terabyte jobs.
Anyone building or operating large‑scale data pipelines should understand MapReduce’s model and design choices, as they underpin Hadoop, Spark, and many modern processing systems.
9/10

