# monitor.mttr
The reliability figures a status report or an SRE review quotes for a
window: how many incidents, how much downtime, mean time to recovery (MTTR)
and mean time between failures (MTBF). Feed it `monitor.incidents`' output.
## Definitions
The window is `from <= t < to`.
- **An incident counts** when it overlaps the window: it starts before `to`
and either starts at or after `from` or is still going after `from`. One
that ends exactly at `from`, or starts exactly at `to`, belongs to the
neighbouring window. An ongoing incident (`end` null) lasts
`durationSeconds`, which `monitor.incidents` measured up to its `now`.
- **downtimeSeconds** is incident time inside the window only, so back-to-back
windows never count the same second twice.
- **mttrSeconds** is the mean *full* duration of the resolved incidents that
count, rounded half up; null when none has resolved. Full, not clipped: an
outage that began before the window still took that long to fix. Ongoing
incidents are left out because their repair time is not known yet. This is
the usual "total resolution time / number of incidents" definition of mean
time to recovery.
- **mtbfSeconds** is the operational time in the window divided by the
number of incidents, `(to - from - downtimeSeconds) / incidents`, rounded
half up; null with no incidents (no failures, no mean between them). This
is the reliability-engineering definition: "the sum of the lengths of the
operational periods divided by the number of observed failures". Some
tools instead measure MTBF from one failure's start to the next's, which
includes the repair time; this one does not.
- **longestSeconds** is the longest full duration among the counted
incidents, 0 when none.
Worked example: incidents of 600 s and 300 s in one day (86400 s) give
downtime 900, MTTR 450, MTBF (86400 - 900) / 2 = 42750 s.
## Errors
- `from must not be after to`
- `from and to must be whole seconds`
- `incidents must be in time order and must not overlap` (touching is fine)
- `incident end must not be before its start`
- `durationSeconds must equal end - start` for a resolved incident
- `durationSeconds must not be negative`
## Sources
- Wikipedia, "Mean time between failures", https://en.wikipedia.org/wiki/Mean_time_between_failures
(MTBF = sum of operational periods / number of failures).
- Wikipedia, "Mean time to recovery", https://en.wikipedia.org/wiki/Mean_time_to_recovery
(MTTR = total resolution time / number of incidents).