proomt

Search

Search posts, papers, and topics

All posts

Salesforce EngineeringScott Nyberg5 min readadvanced

How Intelligent Load Shedding Prevents Cascading Failures in Tier-0 Systems

Summary

Salesforce's Cloud Atlas team redesigned service protection for their Tier-0 identity platform to prevent cascading failures from unpredictable traffic. They implemented intelligent load shedding based on queue time and coordinator-free global quota management to ensure five nines availability and tenant fairness.

  • Per-instance rate limiting becomes insufficient with autoscaling and dynamic traffic, leading to outdated limits and inefficient capacity usage.
  • Multi-tenant systems require global quota management to prevent 'noisy neighbors' from consuming shared resources and impacting other tenants.
  • Intelligent load shedding monitors queue time as an early indicator of system pressure, allowing graceful degradation before cascading failures begin.
  • Coordinator-free global quota management enables servers to make independent throttling decisions, avoiding a central bottleneck while ensuring fairness.

Engineers building highly available, multi-tenant distributed systems can learn practical strategies for preventing cascading failures and managing unpredictable traffic in critical services.

7/10

Related reading

  1. How Uber Protects Against Retry Storms

    Uber developed a context-aware mechanism to prevent retry storms in deep microservice dependency chains. It introduces "error ownership" where services claim errors they originate and unclaim errors they propagate, allowing upstream callers to make informed retry decisions and avoid amplifying load on already struggling services.

    Hacker News front pageuber.com12 minHN11949
  2. Architecting Secure Identity: Extending Auth0 PrivateLink with Intelligent Gateways

    Auth0 PrivateLink provides a single private connection, which can be challenging for enterprises needing to route traffic to multiple isolated VPCs like staging and production. The "Intelligent Gateway" pattern solves this by deploying an API Gateway or reverse proxy at the PrivateLink termination point to fan out traffic based on context added by Auth0 Actions.

    Auth0auth0.com3 min
  3. OpenTelemetry everywhere: Migrating a metrics platform at scale

    Atlassian replaced its decade‑old gostatsd‑based metrics pipeline with a fully OpenTelemetry‑based stack by keeping the StatsD‑UDP contract on the client side and swapping in purpose‑built OTel Collector distributions for collection, ingest, aggregation, and forwarding. The migration was done incrementally, saved ~3.9% CPU per service, cut sidecar cost ~30% fleet‑wide, halved aggregation CPU, and…

    CNCFcncf.io6 minHN1
  4. Metastable Failures in Distributed Systems

    This paper introduces and formalizes "metastable failures" in distributed systems, a class of outages where a trigger pushes a system into a bad state that persists due to a sustaining effect, even after the trigger is removed. These failures often stem from features designed for efficiency or reliability and require significant external intervention to resolve.

    Hall of Famesigops.org25 minpaperHN16112