InfoQMark Silvester3 min readintermediate
Lyft Moves Streaming Fleet to Apache Flink Kubernetes Operator
Summary
Lyft migrated its hundreds of production Flink jobs from a home‑grown Kubernetes operator to the Apache Flink Kubernetes Operator, gaining last‑state upgrades, in‑place autoscaling, and resource autotuning while cutting typical deployment downtime to 3–6 minutes. The switch also enabled Flink 1.19 features, saved millions in over‑provisioned capacity, and required workload‑specific scaling strate…
- Moving to the Apache Flink Kubernetes Operator enabled last-state upgrades, in-place autoscaling, and resource autotuning, reducing typical deployment downtime to 3–6 minutes (20 minutes for the largest jobs).
- The operator's built-in last-state mode restores from HA metadata or latest checkpoint, avoiding the legacy operator's savepoint retry and idempotency problems.
- Upgrading to Flink 1.19 unlocked in-place scaling and KinesisStreamsSource metrics, allowing the autoscaler to right-size the fleet and save millions of dollars per year.
- Autotuning that resizes container memory requires pod restarts, so Lyft split workloads: critical jobs use in-place autoscaling, less critical jobs accept restarts for autotuning.
Engineers running streaming workloads on Kubernetes need to understand the operational benefits and trade‑offs of adopting the Apache Flink operator versus a custom solution.
6/10



