Functional Weave
Code in Rust

monitor.alert-state@1.0.0

README.md

3,115 bytes · view raw

# monitor.alert-state

The lifecycle of one alert, one evaluation at a time, and the notification
each step should send. The function is pure: store the returned `state` and
pass it back as `previous` next time. Start from
`{ phase: "inactive", since: null, lastNotifiedAt: null, clearSince: null }`.
The condition itself comes from anywhere: `monitor.alert-rule`, a heartbeat,
a burn-rate alert.

## Transitions

| previous | condition | result |
| --- | --- | --- |
| inactive or resolved | false | unchanged |
| inactive or resolved | true | `pending` since now; or, when `forSeconds` is 0, `firing` since now and notify `firing` |
| pending | true | `firing` since now and notify `firing` once `now - since >= forSeconds`, else unchanged |
| pending | false | `inactive` since now, no notification (it never fired, so there is nothing to resolve) |
| firing | true | stays firing, `clearSince` reset to null; notify `repeat` when `renotifySeconds` is set and `now - lastNotifiedAt >= renotifySeconds` |
| firing | false | `clearSince` is set to now if it was null; once `now - clearSince >= clearForSeconds` it becomes `resolved` since now and notifies `resolved` |

- `lastNotifiedAt` is the time of the last notification of any kind, and is
  carried through the other phases; entering `firing` sets it, so the repeat
  interval counts from the first notification.
- `changed` says whether the **phase** changed, so a caller knows when to
  write an alert history row. Store the returned state every time anyway:
  `clearSince` and `lastNotifiedAt` can change without the phase changing.
- `since` of a firing alert is when it started firing, not when it became
  pending.

## Why this shape

`forSeconds` is Prometheus's `for`: "wait for a certain duration between first
encountering a new expression output vector element and counting an alert as
firing"; until then the alert is pending, and a pending alert whose condition
goes away is simply dropped. `clearForSeconds` is the other direction, what
Prometheus calls `keep_firing_for` ("keep this alert firing for the specified
duration after the firing condition was last met"): a check that flaps
true/false/true sends one firing and one resolved, not a storm. Here the
condition must stay false for the whole delay; one true evaluation during it
cancels the resolve. `renotifySeconds` matches Alertmanager's
`repeat_interval`: a reminder while an alert keeps firing.

## Errors

- `now 900 is earlier than since 1000` (also for `lastNotifiedAt` and
  `clearSince`): time never goes backwards between evaluations.
- `unknown alert phase: X`
- `since must be set in phase pending` (or `firing`)
- `forSeconds must not be negative`, `clearForSeconds must not be negative`
- `renotifySeconds must be null or at least 1`
- `now must be a whole number of seconds`

## Sources

- Prometheus, "Alerting rules", https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/
  (`for`, pending and firing, `keep_firing_for`).
- Prometheus Alertmanager, "Configuration", https://prometheus.io/docs/alerting/latest/configuration/
  (`repeat_interval`).