Functional Weave
Code in Rust

monitor.alert-rule@1.0.0

README.md

3,129 bytes · view raw

# monitor.alert-rule

Evaluates one threshold alert rule ("CPU above 90% for 5 minutes", "free
disk below 10 GB") over a stored metric series at a moment `now`, and says
whether it is `inactive`, `pending`, `firing` or `no-data`.

## How it decides

1. Only samples with `at <= now` are looked at (a series stored ahead of an
   evaluation time, as in a replay, does not leak the future). The whole
   series must still be in strictly ascending time order.
2. The samples are walked in order, carrying a breach:
   - not breaching: a sample *beyond* `threshold` (by `op`: `gt` is
     `value > threshold`, `gte` `>=`, `lt` `<`, `lte` `<=`) starts a breach;
     `since` is that sample's `at`.
   - breaching: the breach continues while the value is still beyond
     `clearThreshold` by the same comparison, and clears on the first sample
     that is not. With `clearThreshold` null it is the threshold itself, so
     there is no hysteresis.
3. No sample at or before `now`: **no-data**. A `staleAfterSeconds` limit and
   a latest sample more than that old (`now - at > staleAfterSeconds`):
   **no-data** too, because a silent exporter is not a healthy one; `value`
   then reports that stale latest value, and `since` is null.
4. Latest sample breaching: **firing** when `now - since >= forSeconds`,
   otherwise **pending**. Not breaching: **inactive**.

`heldSeconds` is `now - since` while breaching (pending or firing), else 0.

## Why this shape

The `for` duration follows Prometheus alerting rules: "The optional `for`
clause causes Prometheus to wait for a certain duration between first
encountering a new expression output vector element and counting an alert as
firing for this element", and "Elements that are active, but not firing yet,
are in the pending state." Prometheus fires once the alert has been active for
at least `for` (`>=`), and with no `for` it fires on the first evaluation;
`forSeconds: 0` does the same here.

Hysteresis (a separate clear level) is what stops an alert from flapping when
a metric hovers around its threshold: "above 90 to fire, back under 80 to
clear". `clearThreshold` must be on the non-breaching side: at or below the
threshold for `gt`/`gte`, at or above it for `lt`/`lte`.

A new breach after a clear restarts `since`: the for-duration is about one
continuous breach, not the sum of several. Prometheus behaves the same way:
an alert whose expression stops returning goes back to inactive and starts
over.

Unlike Prometheus this looks at the stored samples, not at a query result per
evaluation, so the breach is judged sample by sample; between samples the
last value holds.

## Errors

- `samples must be in strictly ascending time order` (from monitor.series-window)
- `clearThreshold must be on the non-breaching side of threshold: gt 90, clear 95`
- `forSeconds must not be negative`
- `staleAfterSeconds must not be negative`
- `unknown comparison: eq`
- `now must be a whole number of seconds`

## Sources

- Prometheus, "Alerting rules", https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/
  (the `for` clause, pending and firing states).