FundamentalsIntermediate

Incident Management: Framework, Process & Tools

Incident management explained: the IT incident management framework, process steps, severity levels, roles (incident commander, scribe), postmortems, and the tools that make modern incident response work.

11 min read
Atatus Team
Updated October 1, 2026
8 sections
01

What is incident management?

More than paging someone when things break

Incident management is the process of detecting, responding to, and learning from unplanned disruptions to a service. The goal: minimize impact to users and reduce the probability of recurrence.

IT incident management originated in ITIL (Information Technology Infrastructure Library). Modern SRE and DevOps teams have evolved it into a lighter, more automation-driven discipline.

An incident is not every bug or ticket. By convention, an incident is a service disruption that affects users or risks doing so. Severity levels (SEV-1 through SEV-5) grade impact.

Incident management is distinct from problem management (fixing the underlying cause once an incident is resolved) and change management (preventing changes from causing incidents).

02

The incident management framework

Six phases of a well-run incident

Detect: automated monitoring or human report of a disruption. The faster detection, the lower MTTD and overall MTTR.

Triage: assess severity, declare the incident, assign roles. Severity is a decision — SEV-1 (critical: user-visible outage), SEV-2 (major: partial impact), SEV-3 (minor: degraded performance), SEV-4 (minor: cosmetic).

Respond: incident commander (IC) coordinates; subject-matter experts investigate; scribe documents. Communication channels (dedicated Slack channel, status page) activated.

Mitigate: restore service — often via rollback, failover, or runbook action. Mitigation is NOT fix; the goal is to stop user impact fast.

Resolve: close the incident once service is restored and the mitigation is stable. Update stakeholders, close status page incident.

Learn: run a blameless postmortem within 48–72 hours. Document timeline, root cause, impact, and action items.

03

Incident roles and responsibilities

Who does what during a SEV-1

Incident Commander (IC): coordinates the response. Not an engineer — a decision-maker. Decides what to try next, who to page, when to update stakeholders.

Communications Lead: handles external updates — status page, customer email, exec briefings. Keeps the IC free to focus on technical response.

Scribe: documents the timeline in real-time. Timestamps of alerts, decisions, actions, outcomes. Critical for the postmortem.

Subject-Matter Experts (SMEs): engineers with context on the affected systems. Rotate in and out as the investigation narrows.

For small teams, one person may wear multiple hats. The important thing is that the roles exist, not that they are different people.

04

Severity levels

How to classify incidents consistently

SEV-1 (Critical): service is down or severely degraded for most users. All hands on deck. Status page red. CEO may need to be informed.

SEV-2 (Major): significant impact for a subset of users, or risk of becoming SEV-1. Full response team activated. Status page yellow.

SEV-3 (Minor): limited impact, degraded performance, workaround exists. On-call engineer + 1 SME handle during business hours.

SEV-4 (Trivial): cosmetic issues, non-critical alerts, early-warning signals. Ticketed, no urgent response.

Document severity criteria in your incident runbook. Debating severity during an incident wastes response time.

05

Blameless postmortems

The ritual that turns incidents into reliability improvements

Blameless means focusing on systems and processes, not individuals. "Why did the deploy pipeline let a broken build reach prod?" instead of "Why did Alice push bad code?".

Timeline: minute-by-minute record of what happened. Who got paged, when; what was tried, what worked, what didn't.

Root cause analysis: 5 Whys, fishbone diagram, or similar. Avoid stopping at the first proximate cause.

Action items: 3–5 concrete, assigned, dated items. Each has an owner and a target date. Follow-through is tracked in sprint planning.

Share the postmortem widely — within engineering at minimum, often across the whole org. Transparency builds the culture.

06

Incident management tools

The modern incident stack

Alerting: PagerDuty, Opsgenie, VictorOps (Splunk), xMatters. Route alerts to the right on-call engineer with escalation.

ChatOps: dedicated Slack channel per incident, with bots that create channels, invite roles, start recording. Common tools: PagerDuty's channel integration, Rootly, Incident.io.

Status pages: Statuspage (by Atlassian), Instatus, Freshstatus. Public communication to customers during incidents.

Incident management platforms: Incident.io, Rootly, Jeli (acquired by PagerDuty), FireHydrant. Integrate alerting + ChatOps + timeline + postmortem workflow.

Observability: Atatus, Datadog, New Relic, Grafana. The investigation tooling — unified APM, logs, metrics shortens MTTR dramatically.

Runbooks: documented in Notion, Confluence, GitHub, or inline in alerts. Automated where safe (service restart, credential rotation).

07

Metrics for incident management

What to track to improve over time

MTTD (Mean Time to Detect): time from issue starting to someone/something noticing. Observability investment reduces MTTD.

MTTA (Mean Time to Acknowledge): time from alert fired to on-call responding. Should be under 5 minutes for SEV-1.

MTTR (Mean Time to Resolve/Repair): total time from start to resolution. The headline reliability metric.

Incident volume by severity: trend over time. Rising SEV-1 count means reliability is degrading.

Postmortem action item completion rate: do we actually follow through? Teams that complete 90%+ of action items see MTTR drop; teams at 30% see it stall.

Oncall burden: pages per rotation. If a rotation gets paged 20 times a week, that is a reliability problem, not a scheduling problem.

08

Common incident management mistakes

What goes wrong even in experienced teams

No incident commander: the technical investigation and the coordination both suffer. Appoint an IC in the first 5 minutes.

Debating in the main channel: side-chats and parallel investigations fragment the response. Keep ONE channel for the incident; branch out only when necessary.

Trying to fix before mitigating: 2 hours of debugging the root cause while users are offline is bad. Mitigate first (rollback, failover), fix after.

Skipping the postmortem: "we were busy, we'll do it next time". Without the ritual, lessons don't compound and the next incident is identical.

Blame culture: when postmortems identify individuals, people hide information in the next incident. Blameless is a cultural requirement, not a nice-to-have.

Key Takeaways

  • Incident management = detect, triage, respond, mitigate, resolve, learn.
  • Severity levels (SEV-1 to SEV-4) determine response urgency and scope.
  • Roles: Incident Commander, Communications Lead, Scribe, SMEs — appointed within 5 minutes.
  • Mitigate before fixing: stop user impact first, root-cause later.
  • Blameless postmortems with 3–5 concrete action items per significant incident.
  • Stack: alerting (PagerDuty) + ChatOps (Slack + Incident.io) + status page + observability (Atatus/Datadog).
  • Metrics: MTTD, MTTA, MTTR, incident volume, action item completion rate.
Get started today

Monitor your applications with Atatus

Put the concepts from this guide into practice. Set up full-stack observability in minutes with no credit card required.

No credit card required14-day free trialSetup in minutes

Related guides