proomt

Search

Search posts, papers, and topics

All posts

CNCFIris Grace Endozo, Farzad Vazirnia and Albert Kerr, Atlassian6 min readadvanced

OpenTelemetry everywhere: Migrating a metrics platform at scale

Summary

Atlassian replaced its decade‑old gostatsd‑based metrics pipeline with a fully OpenTelemetry‑based stack by keeping the StatsD‑UDP contract on the client side and swapping in purpose‑built OTel Collector distributions for collection, ingest, aggregation, and forwarding. The migration was done incrementally, saved ~3.9% CPU per service, cut sidecar cost ~30% fleet‑wide, halved aggregation CPU, and…

  • Preserve the existing client contract (StatsD over UDP) and let the collector speak both StatsD and OTLP to avoid a massive re‑instrumentation effort.
  • Use a single OTel Collector distribution per pipeline stage so you can iterate on one stage without touching the others.
  • Replace the sidecar stack (StatsD + tracing) with a unified collector sidecar, saving ~3.9% CPU per service and ~30% sidecar cost at scale.
  • Introduce a stream‑ID based load‑balancing exporter to spread a single high‑traffic service across shards, eliminating hot‑shard CPU spikes and tightening autoscaling.

The post shows a pragmatic path to migrate a massive, production‑critical metrics pipeline to OpenTelemetry without breaking existing services, delivering measurable cost reductions and operational simplifications that other large‑scale orgs can replicate.

8/10

Related reading

  1. How Meshtastic Metrics Exporter Turned Eight Prometheus Queries Into One with Tiger Data

    A Python exporter reads Meshtastic mesh telemetry from MQTT and writes it directly into a TimescaleDB‑enabled PostgreSQL instance, replacing a Prometheus setup that hit ~0.5 M series for 10 k nodes. The single database lets Grafana dashboards join node metadata with eight time‑series tables in one query, eliminates scrape‑budget limits, and uses native retention/compression (14‑day compression, 3…

    Timescaletigerdata.com6 min
  2. Elastic announces GA of cross-project search on Serverless, enabling teams to query across all linked projects without moving a byte

    Elastic Cloud Serverless now offers General Availability of Cross‑Project Search (CPS), letting you run a single query across up to 100 linked projects (any region, cloud, or solution type) without moving data. Setup is UI‑driven, permissions are evaluated per‑project, and costs are $0.009 / GB retained plus $0.05 / GB egress. CPS works for Observability, Security, Search, vector DB, and ML jobs,…

    Elasticelastic.co6 minrelease
  3. Reimagining service delivery in the agentic era with Google Public Sector

    Google Public Sector’s blog post outlines how several U.S. state and local agencies are using Google Cloud’s AI services (BigQuery, Gemini, document‑analysis models) to replace manual data‑entry pipelines, consolidate data silos, and accelerate specific projects. Reported outcomes include mapping 52 k parcels in <1 yr vs. a 33.5‑yr estimate (UT DOT), $1.3 M cost savings for real‑time translation…

    Google Cloud Bloggoogle.com4 min
  4. Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

    Vercel’s September AI Gateway Production Index shows open‑weight models processing 56% of token volume (up from 7% in Dec 2025) while accounting for only 14% of spend. Token price fell 23.2% month‑over‑month. Anthropic’s Opus 5 captured 22.5% of spend, overtaking Fable 5 which dropped to 4.9%. OpenAI’s new GPT‑6 Astra grabbed ~7.7% of total gateway spend in its first 12 days, more than double Ant…

    Vercelvercel.com6 minHN2
  5. CloudBees vs Harness: Why Migration Isn't the Fix

    The article argues that Harness’s “free migration” offer hides significant downstream costs (training, pipeline rebuilds, compliance recertification) and that even after migration you still lack unified governance across heterogeneous CI/CD tools. CloudBees positions its Unify control plane as a tool‑agnostic layer that adds visibility, continuous governance, AI‑driven test selection, and hybrid…

    Codeshipcloudbees.com5 min
  6. Lyft Moves Streaming Fleet to Apache Flink Kubernetes Operator

    Lyft migrated its hundreds of production Flink jobs from a home‑grown Kubernetes operator to the Apache Flink Kubernetes Operator, gaining last‑state upgrades, in‑place autoscaling, and resource autotuning while cutting typical deployment downtime to 3–6 minutes. The switch also enabled Flink 1.19 features, saved millions in over‑provisioned capacity, and required workload‑specific scaling strate…

    InfoQinfoq.com3 min