FundamentalsBeginner

MTTR Formula: How to Calculate Mean Time to Repair (With Examples)

Learn the MTTR formula, how to calculate Mean Time to Repair, see worked examples, and understand the related MTTF, MTBF, and MTTD metrics for incident response and reliability engineering.

10 min read
Atatus Team
Updated October 1, 2026
6 sections
01

The MTTR formula

The simple arithmetic, before the nuance

MTTR (Mean Time to Repair) is the average time it takes to restore service after a failure. The formula is: MTTR = Total downtime caused by incidents ÷ Number of incidents.

Example: If your service had 5 incidents last quarter that caused a combined 10 hours of downtime, MTTR = 10 hours ÷ 5 = 2 hours.

MTTR is measured from the moment a failure begins affecting users to the moment service is fully restored. Partial restoration (e.g., only one region recovered) does not stop the clock.

The industry convention is to express MTTR in minutes for high-availability services and hours for less critical systems. Elite DevOps teams target under 1 hour for business-critical services (DORA 2023 report).

02

Worked example: calculating MTTR for a web service

A real-world calculation with multiple incidents

Over 30 days a payments API experienced 4 incidents: Incident 1 lasted 15 minutes (database connection pool exhaustion), Incident 2 lasted 45 minutes (expired SSL certificate), Incident 3 lasted 8 minutes (bad deployment, auto-rollback), and Incident 4 lasted 92 minutes (upstream DNS provider outage).

Total downtime = 15 + 45 + 8 + 92 = 160 minutes. Number of incidents = 4. MTTR = 160 ÷ 4 = 40 minutes.

Observe that one long incident (92 min) pulls the mean upward. Many teams supplement MTTR with the median (middle value) to reduce the impact of outliers. In this example the median is (15 + 45) ÷ 2 = 30 minutes.

Weighted MTTR counts each minute of downtime equally regardless of incident count. A single 4-hour outage and twelve 20-minute outages have the same total downtime (240 min) and the same business impact — but the first has MTTR = 240 min, the second has MTTR = 20 min. Pair MTTR with total downtime for a complete picture.

03

MTTR vs MTTF vs MTBF vs MTTD

The four reliability metrics that get confused

MTTF (Mean Time To Failure): average time a non-repairable component runs before failing. Used for hardware (disks, batteries) that are discarded on failure, not fixed.

MTBF (Mean Time Between Failures): average time between successive failures of a repairable system. MTBF = MTTF + MTTR for the same component. Higher MTBF = more reliable system.

MTTD (Mean Time To Detect): average time from the start of an incident to the moment someone (or an alert) notices it. MTTD is a leading indicator — if detection is slow, repair is slow.

MTTR (Mean Time To Repair): the one we calculated above — detection + diagnosis + resolution combined. Sometimes split into MTTR = MTTD + MTTI (investigate) + MTTF (fix).

The relationship: Total incident time = MTTD + MTTI + MTTR-fix. Optimizing any of these reduces user-visible downtime. Observability platforms directly reduce MTTD and MTTI; runbook automation and SRE practices reduce MTTR-fix.

04

How to reduce MTTR

Practical steps that cut MTTR in half (or better)

Reduce detection time (MTTD) with proper alerting. Alert on symptoms users experience (error rate, latency), not on causes (CPU, memory). Alerts should fire within 1 minute of SLO violation, not when a disk fills up at 3am and nothing is actively broken yet.

Reduce diagnosis time with distributed tracing and correlated logs. If finding the root cause requires grep across 20 log files on 50 hosts, you have a tooling problem, not a people problem. OpenTelemetry + unified observability (Atatus, Datadog, etc.) cuts diagnosis time by 60–80%.

Automate common recovery actions. If 30% of your incidents are "restart the service", put that in a runbook with a one-click action. If 20% are "rotate the credential", automate it. Automation outperforms humans in the small hours of the morning.

Practice incident response. Game days (planned incident simulations) uncover gaps in playbooks, oncall rotations, and tooling before a real incident finds them. Netflix's Chaos Monkey program exists precisely to make real incidents routine.

Blameless postmortems after every significant incident. Each postmortem produces 2–3 action items that reduce the probability of recurrence or the time to recover. Over a year this compounds dramatically.

05

MTTR benchmarks by industry

What good looks like, by DORA class

DORA 2023 "Elite" performers: MTTR under 1 hour. "High" performers: 1 day to 1 week. "Medium": 1 week to 1 month. "Low": over 1 month.

E-commerce and payments: elite teams target MTTR < 30 minutes for revenue-affecting services. Every minute of outage costs real money — $300k/hour was the Amazon figure in 2013; higher now.

SaaS B2B: MTTR < 2 hours is common for business-hours services. SLAs typically guarantee 4-hour recovery with credits for violations.

Internal tools and non-revenue services: MTTR < 1 business day is acceptable. The tradeoff is lower oncall burden vs faster recovery.

What matters more than absolute MTTR is the trend. If MTTR is dropping quarter over quarter, your reliability engineering is working. If it is flat or rising, you need to invest in observability or process.

06

Measuring MTTR correctly: common mistakes

Pitfalls that make your MTTR number meaningless

Including planned maintenance: scheduled downtime for a database upgrade is not an incident. Keep MTTR for unplanned outages only, or you cannot compare across quarters.

Starting the clock when the ticket was filed, not when the user impact began: if a monitoring gap means detection took 20 minutes, that time counts toward MTTR, not toward a happier number.

Rounding to the nearest hour: 23-minute and 59-minute incidents both get recorded as "1 hour". You lose signal on the fast fixes that your team actually did well.

Treating all incidents equally: a 5-minute Severity-1 (homepage down) and a 5-minute Severity-3 (admin dashboard slow) have very different business impact but the same MTTR contribution. Report MTTR broken out by severity.

Not measuring at all: "we feel our incidents are getting shorter" is not data. Capture incident start, detection, acknowledgment, resolution timestamps automatically from your alerting + incident management tooling (PagerDuty, Opsgenie, Incident.io).

Key Takeaways

  • MTTR = Total incident downtime ÷ Number of incidents.
  • MTTR is distinct from MTTF (component lifetime), MTBF (between failures), and MTTD (time to detect).
  • Reduce MTTR by cutting detection time (better alerts), diagnosis time (distributed tracing), and resolution time (automation + runbooks).
  • Elite DevOps teams (DORA) target MTTR under 1 hour; measure trend over time, not just absolute values.
  • Pair MTTR with total downtime and median; MTTR alone hides outliers.
  • Observability platforms with unified APM, logs, and traces are the single biggest MTTR lever.
Get started today

Monitor your applications with Atatus

Put the concepts from this guide into practice. Set up full-stack observability in minutes with no credit card required.

No credit card required14-day free trialSetup in minutes

Related guides