What is log analysis?
Turning log noise into operational signal
Log analysis is the process of parsing, searching, aggregating, and interpreting log data to answer operational, security, and compliance questions.
Logs are the oldest telemetry source — every application and system emits them — but raw logs are low signal: hundreds of GB per day of mostly-redundant text.
Modern log analysis combines structured logging (JSON log entries with named fields), centralized ingest (ship from every host to one place), full-text search, and ad-hoc analytics (aggregations across billions of events).
The two main workloads: operational (why did this request fail?) and security (did we have unauthorized access last Tuesday?). Both benefit from the same storage; the UI often differs.
The modern log analytics stack
Components of a production log pipeline
Agent: a lightweight shipper on each host that tails files or reads stdout. Filebeat, Fluent Bit, Vector, or OpenTelemetry Collector.
Transport: often a buffer (Kafka, Redis Streams) between agents and ingest to absorb spikes.
Ingest + storage: Elasticsearch, OpenSearch, ClickHouse, Loki, or a managed platform (Atatus, Datadog, Splunk).
Query layer: Lucene query syntax (Elasticsearch), SPL (Splunk), LogQL (Loki), SQL (ClickHouse). The query language shapes the whole workflow.
UI: Kibana, Grafana, Splunk Web, or a managed vendor UI. Investigation UX matters more than feature checklists.
Alerting: trigger on query results exceeding thresholds. Elastic Alerting, Grafana Alerting, Splunk Alerts.
Structured logging: the single highest-leverage change
Why plain-text logs are a tax
Plain-text logs force every reader to re-parse them. Grep across production logs for "user_id=42" fails if one service writes userId=42 instead.
Structured logging emits logs as JSON (or similar) with named fields: {"timestamp":"...", "level":"error", "service":"api", "user_id":42, "trace_id":"...", "message":"..."}.
Benefits: fields are indexed and queryable. "Find all errors for user_id=42 in the last hour" is a one-field filter, not a regex nightmare.
Adopt a logging library that supports structured output: Winston or Pino (Node.js), Zap or Zerolog (Go), Serilog (.NET), structlog (Python), Monolog with JSON formatter (PHP).
Enforce a schema: every log entry carries timestamp, service, env, level, trace_id, message, and context-specific fields.
Log parsing and grok
For when you can't change the log format
Not all logs are under your control. Nginx access logs, syslog, framework defaults — all plain text with implicit structure.
Grok (Logstash + Elasticsearch ingest pipelines) applies named regex patterns to extract fields from unstructured logs.
Example: %{IPORHOST:client_ip} %{USER:ident} %{USER:auth} \[%{HTTPDATE:timestamp}\] "%{WORD:method} %{URIPATHPARAM:path} HTTP/%{NUMBER:http_version}" %{NUMBER:status} %{NUMBER:bytes}
Test grok patterns in Kibana's Grok Debugger or on grokdebugger.com before deploying to production pipelines.
Modern alternative: elastic ingest pipelines with built-in processors (geoip, user_agent, csv, kv) replace much of what grok used to do.
Correlation: logs + traces + metrics
The step that collapses MTTR
Logs alone tell you what happened. Traces alone tell you how a request flowed. Metrics alone tell you how much. Correlating all three is where MTTR drops.
The connector is trace_id. Every log entry emitted during a request includes the trace_id of the originating request. Every APM trace exposes the same trace_id.
In an incident: a metric alert fires → open the service dashboard → click an example trace → see the log lines emitted during that trace → identify the root cause.
OpenTelemetry's log signal (OTel Logs) standardizes this correlation across vendors. Older setups rely on manual trace_id injection.
Unified observability platforms (Atatus, Datadog, New Relic) build this correlation into the UI; separate stacks (ELK + Prometheus + Jaeger) require manual stitching.
Advanced log analysis techniques
Beyond grep
Log patterns and clustering: group similar log lines into templates (e.g., Elastic's Categorize Text aggregation). Finds dominant log patterns in minutes instead of hours of eyeballing.
Anomaly detection: ML models that learn normal log volumes and alert on deviations. Elastic ML, Datadog Watchdog, Splunk ITSI. Noisy by default; tune aggressively.
Rare-term detection: log entries that appear once in a day of production. Often the first indicator of a new error or attack.
Session reconstruction: group logs by session_id or user_id across services for a complete view of one user's journey.
SIEM-style correlation: multi-event rules like "5 failed logins within 1 minute, from the same IP, followed by a successful login". Essential for security log analysis.
Log analysis best practices
Eight lessons from running log analytics at scale
Standardize log schema: every service emits the same base fields (timestamp, service, env, level, trace_id). Reserved field names, no exceptions.
Choose log levels deliberately: INFO is for high-level operation flow, DEBUG for diagnostic detail (off in production), WARN for recoverable issues, ERROR for failures.
Sampling for high-volume services: 100% ERROR, 100% WARN, 10% INFO, 1% DEBUG. Preserves signal while cutting cost.
Retention tiers: hot (7 days, fast search), warm (30 days, cheaper), cold (90+ days, S3-backed). Match retention to query frequency.
Mask PII at the agent: email, SSN, credit card numbers never leave the host raw. Compliance and privacy depend on it.
Don't log secrets: API keys, tokens, passwords — ever. Audit CI pipelines to catch accidental leaks.
Alerts on log patterns, not just thresholds: "any ERROR in payment-service" is more actionable than "ERROR rate above baseline".
Review log usage quarterly: delete indexes nobody queries, archive what you must retain for compliance.
Key Takeaways
- Log analysis = parsing, searching, aggregating, and interpreting logs for operational, security, and compliance insight.
- Modern stack: agent + transport + ingest/storage + query layer + UI + alerting.
- Structured logging (JSON with named fields) is the single highest-leverage change for log usefulness.
- Grok and ingest pipeline processors extract fields from unstructured logs.
- Correlation with traces and metrics via trace_id is where MTTR drops.
- Advanced techniques: clustering, anomaly detection, rare-term detection, SIEM correlation.
- Best practices: standard schema, sampling, retention tiers, PII masking, no secrets.