Saturday, October 3, 2026 · Week 40 DE · EN · FR · ES Dark
News

Atlassian moves 100,000 hosts to OpenTelemetry

Atlassian is moving its metrics pipeline to OpenTelemetry without services changing their code. The aggregation now needs only half the CPU.

By Benedikt Langer September 20, 2026 5 min read
Atlassian moves 100,000 hosts to OpenTelemetry

Atlassian is moving the metrics pipeline for around 100,000 hosts from gostatsd to the OpenTelemetry Collector without a single service having to change its interface. For the same traffic volume, the new aggregation stage needs only about half the CPU. The old gostatsd aggregators and the nomad proxy will keep running for now.

Key takeaways

  • The StatsD interface remains in place, with the OpenTelemetry Collector working behind it. Services send over UDP to the familiar address, while the new Collector distribution also accepts OTLP.
  • A new shard key distributes large services evenly across the pool. The streamID key in the load balancing exporter keeps every time series on the same shard.
  • The aggregation rolls up around 4.8 billion data points per minute to around 220 million. For the same traffic volume, the stage needs only about half the CPU.

Limits of the Previous Pipeline

Atlassian’s metrics pipeline collects telemetry from around 100,000 hosts across 14 regions. Until now it ran on gostatsd, a company-maintained open source implementation of the StatsD protocol.

The pipeline operated with an SLO of 99.95 percent, an internal service level objective. In addition, gostatsd accepted metrics only via UDP and had no solution for traces or logs. The team would have had to rebuild every innovation from the OpenTelemetry community by hand.

What is the OpenTelemetry Collector? The OpenTelemetry Collector is modular software for receiving, processing and forwarding telemetry data. It accepts protocols such as OTLP or StatsD and exports to backends like SignalFx or S3. Retries, queues and throttling are part of the standard feature set, and the OpenTelemetry community develops many of its components collaboratively.

A Rebuild in Four Stages

Atlassian deployed a separate Collector distribution at each of the four stages – collection, ingest, aggregation and forwarding – allowing each stage to be rebuilt individually. Nothing changed at the interface to the applications. The services continue to send StatsD metrics over UDP to the same address, while the new Collector distribution also accepts OTLP, the native protocol of OpenTelemetry.

Even just merging the StatsD and tracing sidecars, the helper processes on every host, brought savings. For the most expensive services, it saved on average about 3.9 percent CPU per service; at fleet scale, sidecar costs fell by around 30 percent.

Even Load Distribution

An aggregator must see every data point of a time series in order to aggregate them. The previous in-house proxy called nomad ensured this by hashing the pair of service and environment to a shard, a fixed slice of the aggregator pool. Because the metric load concentrates on a few large services, their shards came under sustained load.

In the new setup, the load balancing exporter of the OpenTelemetry Collector takes over this task. With the routing key streamID, a hash of all attributes of a time series, it spreads a large service evenly across the pool while every time series stays on the same shard. CPU load per shard shifted from sharp spikes and idle replicas to a flat distribution. Since then, the pool can be scaled down outside peak times.

Aggregation at Half the Effort

The aggregation stage receives around 4.8 billion data points per minute and passes on around 220 million after aggregation. Most metrics at Atlassian come as deltas, meaning changes since the last report. No existing component summarized such deltas the way users expect.

Atlassian therefore wrote the Atlassian Aggregation Processor and released it as open source. It builds on a standard component of the OpenTelemetry community, summarizes delta values every 60 seconds by default and passes other metric types through unchanged.

For the same traffic volume, the new aggregation stage needs only about half the CPU. The aggregators no longer have to parse the gostatsd format, the load is distributed more evenly, and the team inherits optimizations from the community.

Forwarding and the Next Step

Forwarding was previously handled by a self-built internal forwarder. Today it runs as a stateless Collector distribution called metrics-gateway, which distributes the data to SignalFx, S3 and other backends. The Collector ships with retries, queues and throttling under overload.

For serverless functions where no sidecar can run, a self-built OTel Lambda extension replaces the gostatsd extension. It keeps the StatsD address and the environment variables, so code changes are unnecessary.

The old gostatsd aggregators and nomad will keep running for now. Together they still account for around 38 percent of CPU requests in the metrics clusters, with nomad alone at around 13 percent of total resources. Once they are phased out, the pipeline will run on OpenTelemetry end to end.

Iris Grace Endozo, Farzad Vazirnia and Albert Kerr of Atlassian see the next step as moving instrumentation to the OpenTelemetry SDK. So far the services have been instrumented with client libraries from vendors and in-house development, that is, equipped with measurement points in the code, including Datadog’s DogStatsD and StatsD libraries.

Frequently Asked Questions

What is gostatsd at Atlassian?

gostatsd is an open source implementation of the StatsD protocol maintained by Atlassian. It accepts metrics only via UDP and does not cover traces or logs. Together with the in-house proxy nomad, the gostatsd aggregators still account for around 38 percent of CPU requests in the metrics clusters.

Why was the shard key a problem?

The previous proxy hashed the pair of service and environment to a shard. Because the load concentrates on a few large services, precisely those shards came under sustained load. The streamID key spreads a service evenly across the pool and keeps every time series on the same shard.

What changes for service teams?

Nothing changes in day-to-day operation; the services keep sending StatsD over UDP to the same address. For serverless functions, a Lambda extension with the same address and environment variables takes over this role. As a next step, Atlassian plans to move instrumentation to the OpenTelemetry SDK.

Image source: AI-generated (September 2026)

Translated from the German original using artificial intelligence. The German version is authoritative.

Also available in

FrançaisEspañolDeutsch
MBF Media Newsletter

The monthly briefing for decision-makers

Once a month, the MBF Media Newsletter gathers what matters from cloudmagazin, MyBusinessFuture, Digital Chiefs and SecurityToday, curated by the editorial team.

Around 23,500 IT and business decision-makers read this newsletter. Read along.

Subscribe for free
MBF Media Newsletter, aktuelle Ausgabe auf dem iPhone
A magazine by Evernine Media GmbH