proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameJeff Hodges201317 min readintro

Notes on Distributed Systems for Young Bloods

Summary

This article outlines essential lessons for new distributed systems engineers, emphasizing that systems fail often and partially. It covers critical design principles like avoiding coordination, implementing backpressure, and designing for partial availability to build robust systems.

  • Distributed systems fail often and partially; design for this reality.
  • Building robust distributed systems is inherently more expensive than single-machine ones.
  • Minimize coordination between machines; horizontal scalability relies on independence.
  • "It's slow" is the hardest problem to debug; distributed tracing is crucial.

New distributed systems engineers should read this for foundational, practical advice on designing resilient systems and avoiding common pitfalls.

7/10

Related reading

  1. A Note on Distributed Computing

    The paper argues that treating remote objects the same as local ones is fundamentally flawed because distributed systems introduce latency, partial failures, and different memory semantics. It outlines a three‑phase development approach that acknowledges distribution concerns early rather than hiding them.

    Hall of Famearchive.org38 minpaper
  2. Metastable Failures in Distributed Systems

    This paper introduces and formalizes "metastable failures" in distributed systems, a class of outages where a trigger pushes a system into a bad state that persists due to a sustaining effect, even after the trigger is removed. These failures often stem from features designed for efficiency or reliability and require significant external intervention to resolve.

    Hall of Famesigops.org25 minpaperHN16112
  3. Worker Backpressure (Part 1)

    Canva added a lightweight, local backpressure loop to its queue worker library that monitors per‑message success/failure, computes a backoff factor against a configurable failure‑rate set‑point, and throttles the worker’s concurrency. In two real incidents the mechanism kept failure rates under 2 % fleet‑wide, limited DLQ growth to a handful of messages, and maintained throughput without manual i…

    Canvacanva.dev10 min
  4. Lessons From Testing Distributed Systems

    This Jepsen blog post announces a retrospective talk on 13 years of testing distributed systems, linking to slides and a video. The post itself contains no detailed lessons or technical content.

    Jepsenjepsen.io1 min
  5. Simple Testing Can Prevent Most Critical Failures

    A study of 198 production failures in distributed systems (Cassandra, HDFS, etc.) found that 92% of catastrophic outages stemmed from incorrect handling of non-fatal errors. Over half of these could have been prevented by simple testing of error handling code, even without deep system understanding.

    Hall of Fameusenix.org58 minpaperHN4
  6. Impossibility of Distributed Consensus with One Faulty Process

    The FLP paper proves that in an asynchronous distributed system, it's impossible to reach consensus if even one process can crash, assuming no synchronized clocks or reliable failure detection. This fundamental impossibility result means any practical consensus protocol must relax one of these assumptions.

    Hall of Famemit.edu20 minpaperHN164