Codeship12 min readintermediate
CloudBees CI Disaster Recovery(DR) proof of concept using Velero
Summary
The post describes a proof‑of‑concept DR setup for CloudBees CI on EKS using Velero with custom patches for cross‑region EBS snapshot replication, achieving 15‑minute RPO and sub‑hour RTO. It also outlines the required backup of $JENKINS_HOME, metadata, DNS switch, and the experimental nature of the patches.
- Custom Velero patches add cross‑region EBS snapshot replication and parallel snapshot creation, enabling low‑RPO restores.
- Backups include $JENKINS_HOME volumes and all Kubernetes metadata; restores recreate StatefulSets but not agent pods.
- Failover procedure involves restoring snapshots, switching Route 53 DNS, and using CloudBees CI's restore‑detect plugin to alert admins.
- RPO of 15 minutes and RTO under an hour were demonstrated with ~100 managed controllers in the demo.
Kubernetes ops teams running CloudBees CI need a concrete, tested DR workflow and realistic performance expectations.
6/10