FundamentalsIntermediate

Distributed Tracing: The Complete Guide (2026)

Distributed tracing explained: how traces, spans, and context propagation work; OpenTelemetry vs Jaeger vs Zipkin; sampling strategies; and how to debug microservices with traces.

12 min read
Atatus Team
Updated October 1, 2026
7 sections
01

What is distributed tracing?

The fundamental building blocks: traces and spans

Distributed tracing is a method of tracking a single request as it propagates through a distributed system — across services, databases, queues, and third-party APIs. Each tracked operation produces a span; a collection of related spans form a trace.

A span represents a single unit of work: a function call, an HTTP request, a database query, a message processing operation. Each span has a start time, a duration, an operation name, attributes (key-value metadata), and events (timestamped logs).

A trace is the tree (or DAG) of spans that make up a complete request. The root span is the entry point (an HTTP request from the user); child spans are called functions, database queries, or downstream service calls.

Each span has a trace_id (shared across the whole trace) and a span_id (unique to this span). Child spans record their parent_span_id, which is how the tree is reconstructed at query time.

A trace tells a story: "User X requested /checkout at 14:03:22. That request called order-service (took 420ms), which called payment-service (took 180ms), which called the Stripe API (took 1900ms — that's where the time went)."

02

Context propagation: how traces survive service hops

The magic that makes distributed tracing work

For a trace to span multiple services, each service must know the trace_id and current span_id when it makes a downstream call. This is called context propagation.

HTTP context propagation uses headers. The W3C Trace Context standard defines traceparent and tracestate headers that every HTTP library should pass through. Example: traceparent: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01.

When Service A calls Service B, Service A's instrumentation adds the headers. Service B's instrumentation reads them and uses the trace_id to create a child span under the same trace.

For asynchronous messaging (Kafka, RabbitMQ, SQS), context is propagated via message headers. OpenTelemetry defines TextMapPropagator abstractions that work across HTTP, gRPC, and messaging transports.

Older protocols (Jaeger's uber-trace-id, B3 for Zipkin) are still common. Modern deployments standardize on W3C Trace Context, with propagators configured to also accept legacy formats during migration.

03

Instrumenting your services

From "zero traces" to "every request traced"

Auto-instrumentation covers 80% of common cases. OpenTelemetry provides auto-instrumentation libraries for most frameworks (Express, Flask, Spring, Rails, Django) that patch HTTP servers, HTTP clients, and database drivers to emit spans automatically.

For Node.js: npm install @opentelemetry/auto-instrumentations-node, then require it as the first import. All inbound HTTP, outbound fetch/axios, and most DB clients are instrumented automatically.

For Java: run the OpenTelemetry Java agent (-javaagent:opentelemetry-javaagent.jar). It bytecode-patches 100+ libraries at startup — zero code changes required.

Add custom spans for domain operations the auto-instrumentation doesn't know about. Example: tracer.start_as_current_span("apply_discount_rules") around a business-logic function makes it a first-class span in the trace.

Add attributes to spans with business context: user_id, tenant_id, order_id, feature_flag_variant. These become searchable dimensions in your trace analytics.

04

Sampling strategies

How to afford distributed tracing at scale

Storing every trace is expensive. A busy service might emit 100,000 spans per second — storing all of them for 30 days is cost-prohibitive.

Head-based sampling: decide at trace start whether to record this trace. Common default: sample 1–10% of traces uniformly at random. Simple to implement; may miss rare error cases.

Tail-based sampling: record all spans, decide at trace end which to keep. Keep 100% of error traces, 100% of slow traces, 5% of normal traces. Requires a stateful collector (OpenTelemetry Collector, Grafana Agent, Vector) that buffers spans until the trace completes.

Rate-limited sampling: record up to N traces per second per service. Prevents cost explosion during traffic spikes but can miss relevant traces during incidents.

Error-biased sampling: always record traces with errors or slow spans; sample normal traces lightly. The industry best practice for production APM.

Managed platforms like Atatus apply tail-based + error-biased sampling by default so you keep the traces that matter without configuring a sampler yourself.

05

Querying and visualizing traces

From data to insight

The flame graph (or waterfall view) is the standard trace visualization. Time runs left-to-right; spans are stacked vertically to show parent-child relationships. The widest spans are the slowest; the deepest spans are the most nested.

Service maps aggregate traces into a graph of services and their call relationships. Edges show request rate, error rate, and p95 latency. Essential for understanding microservice topology.

Trace search by attribute: find all traces where user_id = 12345, or where http.status_code = 500, or where db.statement contains "SELECT * FROM orders". High-cardinality search is what distinguishes traces from metrics.

Trace analytics: aggregate across millions of traces to answer "what's the p99 latency of /checkout broken down by payment_method?". OpenTelemetry Collector + ClickHouse / Honeycomb / Datadog Trace Analytics / Atatus all support this.

Correlate traces with logs by trace_id. When you find a slow trace, one click should show the log lines emitted during that trace. This correlation is the single biggest MTTR lever modern platforms provide.

06

OpenTelemetry, Jaeger, Zipkin: the ecosystem

Which tool does what

OpenTelemetry (OTel): the current standard for instrumentation and the OTLP wire protocol. It replaced OpenTracing (deprecated 2019) and OpenCensus. Vendor-neutral; emits data in a format every modern backend accepts.

Jaeger: an open-source trace backend originally built at Uber. Receives spans (Jaeger thrift, OTLP, Zipkin formats), stores them (Cassandra, Elasticsearch, Badger), and renders flame graphs. CNCF graduated project.

Zipkin: the original open-source distributed tracing system (from Twitter). Still widely deployed. Simpler than Jaeger but less actively developed.

OpenTelemetry Collector: the vendor-neutral agent that receives telemetry (OTLP, Jaeger, Zipkin, Prometheus), processes it (sampling, filtering, enrichment), and exports it to one or more backends. Runs as a sidecar or gateway.

Managed backends: Datadog APM, New Relic Distributed Tracing, Honeycomb, Lightstep, Grafana Tempo, and Atatus all accept OTLP natively. You instrument once and swap backends by changing one config line.

07

Debugging workflows with traces

How traces actually get used during an incident

Pattern 1: "This endpoint is slow". Find traces for the endpoint with high duration; look at the flame graph; identify the longest span. Typical answer: a slow database query or an N+1 pattern.

Pattern 2: "This user hit an error". Filter traces by user_id and error=true; open the trace; find the span with the exception; read the correlated log lines. Reproduces the exact failure path.

Pattern 3: "Deployment caused a regression". Compare p95 latency of a service before/after the deploy timestamp. Drill into traces to see which newly-slow span is responsible. Common finding: a new SQL query added by the deploy.

Pattern 4: "Service dependency map broken". Open the service map, find the service with elevated error rate downstream of yours, open an example error trace. Common finding: external API returning 429 or 503.

The common thread: traces turn vague user reports into specific, actionable spans. The learning curve is 1–2 weeks of practice during real incidents.

Key Takeaways

  • A trace is a tree of spans tracking one request across services; context propagation via headers is what makes it work.
  • Instrument with OpenTelemetry — vendor-neutral and the industry standard since 2019.
  • Sampling matters at scale: use tail-based + error-biased sampling to keep what's valuable without storing everything.
  • Flame graphs for individual traces; service maps for topology; trace analytics for aggregate answers.
  • Correlate traces with logs by trace_id — this is where MTTR drops dramatically.
  • OTLP is the wire protocol; Jaeger, Zipkin, and managed APM backends all accept it.
Get started today

Monitor your applications with Atatus

Put the concepts from this guide into practice. Set up full-stack observability in minutes with no credit card required.

No credit card required14-day free trialSetup in minutes

Related guides