MTTR, MTTF and the Availability Equation
Two metrics, one formula, one insight: reducing MTTR is almost always more effective than preventing failures.
Your SLO measures availability. That availability is the result of two opposing forces: how often your system fails, and how fast you fix it. Acting on either one has the same mathematical effect — but very different costs.
Definitions
MTTF (Mean Time To Failure): average operating time between two failures. The higher it is, the more reliable your system. MTTR (Mean Time To Recovery): average time from failure onset to return to normal. It includes detection, diagnosis, correction, and verification. MTBF (Mean Time Between Failures) = MTTF + MTTR.
The availability formula
Availability is calculated directly from MTTF and MTTR. It's the ratio of "healthy" time to total time. Example: if your service fails on average every 15 days (MTTF = 21,600 min) and you repair it in 21.6 minutes on average, you achieve exactly 99.9% availability — meaning the full error budget of a 99.9% SLO is consumed.
The MTTR lever is often more effective
Halving MTTR has exactly the same effect on availability as doubling MTTF. Mathematically, M/(M+T/2) = 2M/(2M+T). But improving MTTF means reducing failure frequency — which often requires infrastructure redesign, more testing, chaos engineering. Improving MTTR means writing better runbooks, automating rollbacks, improving dashboards, and training on-call. This is generally faster and cheaper.
Detection time: the hidden part of MTTR
Most teams measure MTTR from alert acknowledgement, not from the actual start of the failure. MTTD (Mean Time To Detect) — the time between failure and first alert — is often invisible in post-mortems. A service that fails at 2am and isn't detected until 8am has a real MTTR of 6h+ even if resolution took 20 minutes. Your SLOs measure real availability, not availability from the alert.
Common pitfalls
- Measuring MTTR from alert acknowledgement: you're underestimating your real MTTR by a factor of 2 to 10 depending on your alerting coverage.
- Treating failures as independent: two services in the same AWS region often fail together — your effective MTTF is much lower than the isolated MTTF.
- Optimising only MTTF: if your MTTR is 4h and your MTTF is 30 days, reducing failures by 20% only gains you 4 minutes of availability per month. Cutting your MTTR to 30 min gains you over 3 hours.
Related articles
Try the simulator →
MTTR/MTTF Optimizer →