Service Level Agreement
SLAMTBF and MTTR
Mean Time Between Failures and Mean Time To Repair — the two fundamental reliability metrics for data centre equipment and systems.
Full Definition
<p>Mean Time Between Failures (MTBF) is the average time between one failure and the next for a repairable system. A higher MTBF indicates greater reliability. MTBF is used to predict component failure rates and to size redundancy requirements.</p><p>Mean Time To Repair (MTTR) is the average time taken to restore a failed component or system to service, including fault diagnosis, logistics and repair or replacement. A lower MTTR indicates a more responsive maintenance operation. Together, MTBF and MTTR determine system availability: Availability = MTBF / (MTBF + MTTR). For a facility targeting 99.9999% availability (Tier IV), both metrics must be very high and very low respectively. Spare parts holding strategies and maintenance contracts are designed around MTBF/MTTR targets.</p>
Also Known As
Source Reference
IEEE Std 493 — Recommended Practice for Reliability Analysis of Industrial and Commercial Power Systems