generator operational EN FR EN
ninesarewelcom
Service objectives, written like a manual.
Menu

User Guide

MTTR, MTTF and the Availability Equation

Two metrics, one formula, one insight: reducing MTTR is almost always more effective than preventing failures.

Your SLO measures availability. That availability is the result of two opposing forces: how often your system fails, and how fast you fix it. Acting on either one has the same mathematical effect — but very different costs.

Definitions

MTTF (Mean Time To Failure): average operating time between two failures. The higher it is, the more reliable your system. MTTR (Mean Time To Recovery): average time from failure onset to return to normal. It includes detection, diagnosis, correction, and verification. MTBF (Mean Time Between Failures) = MTTF + MTTR.

The availability formula

Availability is calculated directly from MTTF and MTTR. It's the ratio of "healthy" time to total time. Example: if your service fails on average every 15 days (MTTF = 21,600 min) and you repair it in 21.6 minutes on average, you achieve exactly 99.9% availability — meaning the full error budget of a 99.9% SLO is consumed.

Availability = MTTF / (MTTF + MTTR)

The MTTR lever is often more effective

Halving MTTR has exactly the same effect on availability as doubling MTTF. Mathematically, M/(M+T/2) = 2M/(2M+T). But improving MTTF means reducing failure frequency — which often requires infrastructure redesign, more testing, chaos engineering. Improving MTTR means writing better runbooks, automating rollbacks, improving dashboards, and training on-call. This is generally faster and cheaper.

Detection time: the hidden part of MTTR

Most teams measure MTTR from alert acknowledgement, not from the actual start of the failure. MTTD (Mean Time To Detect) — the time between failure and first alert — is often invisible in post-mortems. A service that fails at 2am and isn't detected until 8am has a real MTTR of 6h+ even if resolution took 20 minutes. Your SLOs measure real availability, not availability from the alert.

Common pitfalls

  • Measuring MTTR from alert acknowledgement: you're underestimating your real MTTR by a factor of 2 to 10 depending on your alerting coverage.
  • Treating failures as independent: two services in the same AWS region often fail together — your effective MTTF is much lower than the isolated MTTF.
  • Optimising only MTTF: if your MTTR is 4h and your MTTF is 30 days, reducing failures by 20% only gains you 4 minutes of availability per month. Cutting your MTTR to 30 min gains you over 3 hours.

Related articles

Try the simulator →

MTTR/MTTF Optimizer →