Hall of FameAmazon Web Services20174 min readpostmortemadvanced
Summary of the Amazon S3 Service Disruption in the Northern Virginia (US-EAST-1) Region
Summary
An S3 team member accidentally removed too many servers in US-EAST-1 due to an incorrect command input, leading to a multi-hour outage of S3 and dependent AWS services. The incident highlighted challenges in restarting critical subsystems at scale and the need for improved operational safeguards.
- An incorrect command input by an S3 team member removed critical capacity from the index and placement subsystems, causing the outage.
- Restarting core S3 subsystems took longer than expected due to S3's massive growth and infrequent full restarts in large regions.
- The outage impacted S3 GET/LIST/PUT/DELETE requests and dependent AWS services like EC2 launches, EBS, and Lambda.
- Remediation includes modifying operational tools with safeguards, reprioritizing subsystem partitioning, and making the Service Health Dashboard multi-region.
This postmortem offers critical insights for engineers designing and operating large-scale distributed systems on the importance of robust operational tooling, rapid recovery mechanisms, and understanding system dependencies.
8/10


