Functional Weave
Code in Rust

monitor.incidents@1.0.0

README.md

1,863 bytes · view raw

# monitor.incidents

Finds the incidents in a stored series of checks (`monitor.check-status`'s
`Check`): each run of consecutive bad checks is one incident.

## Rules

- **Bad** depends on `incidentStatus`: with `down`, only `down` checks are
  bad (a slow site is not an outage); with `degraded`, `degraded` and `down`
  both are, and one incident can mix them. `up` is refused: an "incident of
  being up" is a configuration mistake.
- A run counts only when it has at least `confirmChecks` checks. This is flap
  suppression, the same idea as "alert after N consecutive failures" in most
  uptime monitors: one lost probe does not make an outage, and would drag
  MTTR down if it did. Shorter runs are dropped entirely, including an
  unconfirmed run still in progress at the last check.
- `start` is the `at` of the first bad check; `end` is the `at` of the first
  check after the run that is not bad (with `incidentStatus: down`, a
  `degraded` check ends a down incident). An incident still bad at the last
  check is ongoing: `end` is null and `durationSeconds` runs to `now`.
- The start is the first check that *saw* the failure, not a guess at when
  it really began between two checks; and the end is the first check that saw
  recovery. So durations are accurate to the check interval, and err long.
- `worst` is `down` if any check in the run was down, else `degraded`.
- `checks` counts the bad checks in the run.

Incidents come out in time order and never overlap, which is what
`monitor.mttr` expects.

## Errors

- `checks must be in strictly ascending time order`
- `unknown check status: X`
- `incidentStatus must be down or degraded, received up`
- `confirmChecks must be at least 1`
- `now 200 is before the last check at 240`: now ends ongoing incidents, so it
  cannot be earlier than the data.
- `now must be a whole number of seconds`