What are log levels?
The severity scale every logging library uses
Log levels are a severity classification applied to each log entry. They let you control how much detail your application emits in different environments — verbose in development, selective in production.
The standard levels, from most verbose to most severe: TRACE, DEBUG, INFO, WARN, ERROR, FATAL.
Setting a log level means "emit all logs at this level or higher". Setting level=INFO suppresses TRACE and DEBUG but emits INFO, WARN, ERROR, FATAL.
Every mainstream logging library (Log4j, Logback, Winston, Pino, Zap, Zerolog, structlog, Monolog) implements the same levels. Minor variations exist (some use NOTICE between INFO and WARN, following syslog).
The six standard log levels
What each means and when to use it
TRACE: extremely fine-grained diagnostic information. "Entered function foo", "Loop iteration 42". Rarely used in production; typically on in local dev only.
DEBUG: diagnostic information useful during development or deep investigation. Variable values, intermediate computation results. Usually off in production.
INFO: high-level operational flow. Service started, request received, batch job completed. The default production level for most applications.
WARN: unexpected conditions that are recoverable or expected under stress. "Retry attempt 2 of 3", "Cache miss, falling back to DB", "Rate limit hit, slowing down".
ERROR: a specific operation failed and user-visible impact may follow. "Payment API returned 500", "Database connection pool exhausted". Always worth attention.
FATAL (or CRITICAL): the application cannot continue. Typically precedes a process shutdown. "Failed to load config, exiting".
Picking the right level
A practical decision tree
Does the log reveal normal operation? → INFO. Service started, request handled, scheduled job ran.
Does the log reveal diagnostic detail only useful to developers? → DEBUG. Variable dumps, step-by-step execution.
Did something unexpected happen, but the system handled it? → WARN. Retry, fallback, degraded response.
Did something fail that users will notice? → ERROR. Request returned 500, external service unavailable.
Did something fail that stops the whole process? → FATAL. Boot config missing, catastrophic data corruption.
Most apps misuse these levels. The common mistake: logging routine events as WARN or every error as FATAL. Reserve higher levels for actionable signal.
Production log level configuration
The settings that actually matter
Default production level: INFO. Captures operational flow without drowning you in debug detail.
Set per-component levels if needed. The HTTP router at WARN+, the database layer at INFO+, the authentication module at DEBUG for a troubleshooting window.
Dynamic log-level changes: libraries like Logback and Winston let you change levels at runtime without restarting. Essential for temporarily cranking up detail during an incident.
Environment-driven config: LOG_LEVEL env var, read at startup. Dev=DEBUG, staging=INFO, prod=INFO.
Avoid shipping DEBUG to centralized log storage in production. Volume costs scale linearly; most DEBUG entries have no value post-incident.
Log levels + sampling
How to control cost without losing signal
For high-volume services, even INFO can be noisy. Sampling retains a percentage of each level: 100% ERROR, 100% WARN, 10% INFO, 1% DEBUG.
Sample at the agent (Filebeat, Fluent Bit, Vector) or at the application level. OpenTelemetry log SDKs support head-sampling.
Tail-sampling (keep all logs for failed traces, sample for successful ones) requires correlation with trace status — supported in modern unified observability platforms.
Never sample FATAL. Never sample security-relevant events (login attempts, access-denied).
Alerting on log levels
From severity to action
Alert on ERROR and above for critical services. Any ERROR in a payment service is worth an engineer's attention.
For noisier services, alert on rate: "ERROR rate 5x baseline for 5 minutes". Reduces noise from one-off errors.
FATAL should trigger a page immediately — the application has stopped working.
WARN typically does not alert; monitor as a trend. Rising WARN rate often precedes rising ERROR rate.
Alert on log pattern, not just level: "any log containing 'segfault' or 'OOM-killed'" catches issues that log levels alone might miss.
Key Takeaways
- Log levels in order: TRACE < DEBUG < INFO < WARN < ERROR < FATAL.
- Production default: INFO. DEBUG and TRACE off in production except during focused investigation.
- INFO = normal flow, WARN = recoverable unexpected, ERROR = user-visible failure, FATAL = cannot continue.
- Set dynamic log levels so you can crank up detail during an incident without a restart.
- Sample INFO and below in high-volume services; never sample FATAL or security events.
- Alert on ERROR+ with rate thresholds; FATAL triggers an immediate page.