Hall of FameAbhishek Verma et al.201560 min readpaperadvanced
Large-scale cluster management at Google with Borg
Summary
Google's Borg is a large-scale cluster manager running hundreds of thousands of jobs across tens of thousands of machines. It achieves high utilization through efficient task-packing and over-commitment, while ensuring high availability and simplifying resource management for users.
- Borg manages heterogeneous workloads (services/batch) across cells of up to tens of thousands of machines.
- High utilization is achieved via admission control, efficient task-packing, over-commitment, and process isolation.
- Jobs are defined declaratively in BCL, supporting rolling updates and fine-grained resource requests.
- Priority bands (monitoring, production, batch) allow preemption of lower-priority tasks for critical workloads.
This foundational paper details the design and operational experience of Google's internal cluster manager, which heavily influenced modern orchestrators like Kubernetes, making it essential for anyone building or operating large-scale distributed systems.
9/10



