FundamentalsBeginner

Traces vs Logs vs Metrics: The Three Pillars of Observability

Traces vs logs vs metrics explained: what each tells you, when to use which, how to correlate them, and why modern observability requires all three.

7 min read
Atatus Team
Updated October 1, 2026
6 sections
01

The three pillars at a glance

A single sentence for each

Metrics tell you how much — numeric time series: request rate, CPU utilization, error count. Cheap to store, cheap to query, low cardinality.

Logs tell you what happened — discrete events with context: "user 42 failed to login at 14:03:22". Higher cardinality, harder to search at scale.

Traces tell you where time went — the end-to-end path of a request across services, with timing for every span. Highest cardinality — each trace is unique.

Modern observability uses all three, correlated by trace_id, in one platform.

02

When to use metrics

The aggregated view

Monitor overall system health: request rate, error rate, p95 latency, apdex, resource utilization.

Set SLOs and alert on them: "alert when availability SLO burn rate exceeds X".

Capacity planning: track trend over weeks and months. Metrics are the only pillar practical for long-term retention of aggregate signal.

Dashboards for broad situational awareness: the status page at a glance.

Weakness: aggregated. "Error rate is up 20%" doesn't tell you which 20% of users or which code path.

03

When to use logs

The discrete event detail

Reconstruct what actually happened for a specific user or request. "What did user 42 do between 2:00 and 2:05?"

Debug errors with full stack traces, parameters, and context that metrics cannot carry.

Audit and compliance: SOC 2, HIPAA, PCI-DSS all require log retention for security events.

Investigate issues retroactively — a trace or metric may not capture an edge case, but a log line will.

Weakness: volume. Logs grow linearly with traffic; indexing cost scales with volume. At scale, retention tiers and sampling are mandatory.

04

When to use traces

The end-to-end request view

Debug latency: "this /checkout request took 2.3s — where did the time go?" The flame graph tells you in one glance.

Microservices root cause: a user-visible error touched 10 services. Which service originated the problem? Traces answer this in seconds.

Service dependency maps: aggregate traces into live topology showing which service calls which, with latency and error at every edge.

High-cardinality analytics: "p99 latency for user tier = premium, region = us-east, feature flag variant = new-checkout". Traces natively support this.

Weakness: cost. Storing every trace is expensive; sampling is required at scale.

05

Correlating the three

The step that makes observability powerful

The correlator is trace_id. Every log line emitted during a request includes the trace_id of the originating request. Every APM trace exposes the same trace_id.

Metrics carry a trace exemplar: a sample trace_id attached to the metric data point, so clicking a high p99 bucket in a chart lands you in a representative trace.

Workflow: metric alert fires → open service dashboard → click a high-latency bucket → see the exemplar trace → see the slow span → see the log lines emitted during that span.

OpenTelemetry's semantic conventions standardize the correlation keys (trace_id, span_id, service.name) across all three signals.

Unified observability platforms (Atatus, Datadog, New Relic, Honeycomb) build this correlation into one UI. Separate stacks (Prometheus + ELK + Jaeger) require manual stitching that slows every incident.

06

The practical investigation workflow

How the three pillars actually get used

1. Alert fires on a metric (SLO burn rate, error rate spike, latency jump).

2. Dashboard shows the shape — which service, which region, which route.

3. Click an example trace exemplar. Flame graph shows the slow or failed span.

4. Click the span to see its log entries. Full context: parameters, stack trace, downstream calls.

5. Identify the root cause and mitigate (rollback, failover, hotfix).

In a well-correlated platform, this is minutes. In a disconnected stack, it is hours.

Key Takeaways

  • Metrics = how much; Logs = what happened; Traces = where time went.
  • Metrics for health + SLOs + alerting; logs for forensic detail; traces for latency and service topology.
  • Correlate via trace_id — the single highest-leverage field across all three signals.
  • OpenTelemetry standardizes the correlation keys and the wire protocol (OTLP).
  • Unified observability platforms collapse the three into one UI; separate stacks fragment the investigation.
Get started today

Monitor your applications with Atatus

Put the concepts from this guide into practice. Set up full-stack observability in minutes with no credit card required.

No credit card required14-day free trialSetup in minutes

Related guides