Functional Weave
Code in Rust

monitor.mttr@1.0.0

README.md

2,520 bytes · view raw

# monitor.mttr

The reliability figures a status report or an SRE review quotes for a
window: how many incidents, how much downtime, mean time to recovery (MTTR)
and mean time between failures (MTBF). Feed it `monitor.incidents`' output.

## Definitions

The window is `from <= t < to`.

- **An incident counts** when it overlaps the window: it starts before `to`
  and either starts at or after `from` or is still going after `from`. One
  that ends exactly at `from`, or starts exactly at `to`, belongs to the
  neighbouring window. An ongoing incident (`end` null) lasts
  `durationSeconds`, which `monitor.incidents` measured up to its `now`.
- **downtimeSeconds** is incident time inside the window only, so back-to-back
  windows never count the same second twice.
- **mttrSeconds** is the mean *full* duration of the resolved incidents that
  count, rounded half up; null when none has resolved. Full, not clipped: an
  outage that began before the window still took that long to fix. Ongoing
  incidents are left out because their repair time is not known yet. This is
  the usual "total resolution time / number of incidents" definition of mean
  time to recovery.
- **mtbfSeconds** is the operational time in the window divided by the
  number of incidents, `(to - from - downtimeSeconds) / incidents`, rounded
  half up; null with no incidents (no failures, no mean between them). This
  is the reliability-engineering definition: "the sum of the lengths of the
  operational periods divided by the number of observed failures". Some
  tools instead measure MTBF from one failure's start to the next's, which
  includes the repair time; this one does not.
- **longestSeconds** is the longest full duration among the counted
  incidents, 0 when none.

Worked example: incidents of 600 s and 300 s in one day (86400 s) give
downtime 900, MTTR 450, MTBF (86400 - 900) / 2 = 42750 s.

## Errors

- `from must not be after to`
- `from and to must be whole seconds`
- `incidents must be in time order and must not overlap` (touching is fine)
- `incident end must not be before its start`
- `durationSeconds must equal end - start` for a resolved incident
- `durationSeconds must not be negative`

## Sources

- Wikipedia, "Mean time between failures", https://en.wikipedia.org/wiki/Mean_time_between_failures
  (MTBF = sum of operational periods / number of failures).
- Wikipedia, "Mean time to recovery", https://en.wikipedia.org/wiki/Mean_time_to_recovery
  (MTTR = total resolution time / number of incidents).