monitor.alert-rule
Evaluate a threshold alert rule over a metric series: inactive, pending, firing or no-data, Prometheus style.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 25 tests, run in TypeScript, Python and Rust.
What it does
Evaluates one threshold alert rule ("CPU above 90% for 5 minutes", "free disk below 10 GB") over a stored metric series at a moment `now`, and says whether it is `inactive`, `pending`, `firing` or `no-data`.
## How it decides
For example
evaluate_alert_rule(samples ×4, op gt, threshold 90, clear threshold —, for seconds 120, stale after seconds —, 180)→ state firing, since 60, value 99, held seconds 120 a breach held for the whole for-duration is firingevaluate_alert_rule(samples ×4, op gt, threshold 90, clear threshold —, for seconds 120, stale after seconds —, 179)→ state pending, since 60, value 97, held seconds 119 one second short of the for-duration is still pendingevaluate_alert_rule(samples ×2, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 60)→ state firing, since 60, value 95, held seconds 0 forSeconds 0 fires on the first breaching sample
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
def evaluate_alert_rule(samples: Sequence[MetricSample], rule: ThresholdRule, now: int) -> RuleEvaluation
| samples | MetricSample[] | in strictly ascending time order; samples after now are ignored |
| rule | ThresholdRule | |
| now | int | Unix seconds of this evaluation |
| returns | RuleEvaluation |
The types it declares, generated into your project
CompareOp = Literal["gt", "gte", "lt", "lte"]
@dataclass(frozen=True)
class ThresholdRule:
"""When a metric is breaching, and for how long before it fires."""
#: how the value is compared with the threshold (gt: value > threshold)
op: CompareOp
#: the value that starts a breach
threshold: int
#: once breaching, the value must pass this the other way to clear; null = threshold
clear_threshold: Optional[int]
#: how long a breach must last before it fires (Prometheus `for`), 0 = at once
for_seconds: int
#: a latest sample older than this is no data; null = never stale
stale_after_seconds: Optional[int]
RuleState = Literal["inactive", "pending", "firing", "no-data"]
@dataclass(frozen=True)
class RuleEvaluation:
"""The rule's state now, and the breach behind it."""
state: RuleState
#: at of the sample that started the current breach; null when not breaching
since: Optional[int]
#: the latest value at or before now; null when there is none
value: Optional[int]
#: now - since while breaching, else 0
held_seconds: int
Your code names it in one line, in the file that uses it
from fune.monitor.alert_rule import evaluate_alert_rule # monitor.alert-rule@^1
Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.
from typing import Optional, Sequence
from .monitor_series_window_types import MetricSample
from .monitor_series_window import series_window ← from monitor.series-window ^1.0.0 · built alongside by fune
from .monitor_alert_rule_types import CompareOp, RuleEvaluation, ThresholdRule
def _beyond(op: CompareOp, value: int, limit: int) -> bool:
if op == "gt":
return value > limit
if op == "gte":
return value >= limit
if op == "lt":
return value < limit
if op == "lte":
return value <= limit
raise ValueError("unknown comparison: %s" % (op,))
def evaluate_alert_rule(samples: Sequence[MetricSample], rule: ThresholdRule, now: int) -> RuleEvaluation:
"""The state of a threshold rule at `now`, Prometheus style: a breach
must hold for `for_seconds` before it fires, and until then it is pending.
With a clear_threshold a breach clears only once the value passes it the
other way, so a metric wobbling around the threshold does not flap
between firing and inactive. The breach clock (`since`) restarts after a
clear."""
if not isinstance(now, int) or isinstance(now, bool):
raise ValueError("now must be a whole number of seconds")
op = rule.op
threshold = rule.threshold
clear = threshold if rule.clear_threshold is None else rule.clear_threshold
_beyond(op, 0, 0)
upward = op in ("gt", "gte")
if (clear > threshold) if upward else (clear < threshold):
raise ValueError("clearThreshold must be on the non-breaching side of threshold: %s %d, clear %d" % (op, threshold, clear))
if rule.for_seconds < 0:
raise ValueError("forSeconds must not be negative, received %d" % (rule.for_seconds,))
stale = rule.stale_after_seconds
if stale is not None and stale < 0:
raise ValueError("staleAfterSeconds must not be negative, received %d" % (stale,))
# series_window checks the order of the whole series and keeps at <= now.
start = min(samples[0].at, now + 1) if len(samples) > 0 else now + 1
seen = series_window(samples, start, now + 1)
since: Optional[int] = None
for s in seen:
if since is None:
if _beyond(op, s.value, threshold):
since = s.at
elif not _beyond(op, s.value, clear):
since = None
if len(seen) == 0:
return RuleEvaluation(state="no-data", since=None, value=None, held_seconds=0)
latest = seen[-1]
if stale is not None and now - latest.at > stale:
return RuleEvaluation(state="no-data", since=None, value=latest.value, held_seconds=0)
if since is None:
return RuleEvaluation(state="inactive", since=None, value=latest.value, held_seconds=0)
held = now - since
return RuleEvaluation(state="firing" if held >= rule.for_seconds else "pending", since=since, value=latest.value, held_seconds=held)Install
fune build
With that line in your source, in a Python project (language python in fune.project), fune build resolves it and its 1 dependency, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:
fune add monitor.alert-rule
The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./monitor.alert-rule-1.0.0-python.fune, or fetch it from a terminal with fune pull monitor.alert-rule@1.0.0:python.
The whole function, every language, is one file too: monitor.alert-rule-1.0.0.fune, 25,566 bytes, sha256 2269259d07d83cc2869899f3d1fabd2898ed5f870612b99233b4bb3593b8a2f3. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
# fune: before monitor.alert-rule
after — your function gets the result and the arguments, and returns the final result.
# fune: after monitor.alert-rule
replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.
# fune: replace monitor.series-window in monitor.alert-rule
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.alert-rule --steps.
# fune: step monitor.alert-rule after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| a breach held for the whole for-duration is firing | samples ×4, op gt, threshold 90, clear threshold —, for seconds 120, stale after seconds —, 180 | → | state firing, since 60, value 99, held seconds 120 |
| one second short of the for-duration is still pending | samples ×4, op gt, threshold 90, clear threshold —, for seconds 120, stale after seconds —, 179 | → | state pending, since 60, value 97, held seconds 119 |
| forSeconds 0 fires on the first breaching sample | samples ×2, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 60 | → | state firing, since 60, value 95, held seconds 0 |
| below the threshold is inactive | samples ×2, op gt, threshold 90, clear threshold —, for seconds 60, stale after seconds —, 60 | → | state inactive, since —, value 80, held seconds 0 |
| gt: a value exactly at the threshold is not a breach | samples ×1, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 0 | → | state inactive, since —, value 90, held seconds 0 |
| gte: a value exactly at the threshold is a breach | samples ×1, op gte, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 0 | → | state firing, since 0, value 90, held seconds 0 |
| hysteresis: dipping below the threshold but above the clear level keeps the breach | samples ×3, op gt, threshold 90, clear threshold 80, for seconds 60, stale after seconds —, 120 | → | state firing, since 0, value 85, held seconds 120 |
| hysteresis: reaching the clear level clears, and a value under the threshold does not re-breach | samples ×3, op gt, threshold 90, clear threshold 80, for seconds 60, stale after seconds —, 120 | → | state inactive, since —, value 85, held seconds 0 |
| a new breach after a clear restarts the clock, it does not reuse the old start | samples ×4, op gt, threshold 90, clear threshold —, for seconds 120, stale after seconds —, 200 | → | state pending, since 120, value 96, held seconds 80 |
| lt with a clear level above: low disk space stays breaching until it recovers past 20 | samples ×3, op lt, threshold 10, clear threshold 20, for seconds 60, stale after seconds —, 150 | → | state firing, since 60, value 12, held seconds 90 |
Show the other 15 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| lt: recovering past the clear level is inactive | samples ×4, op lt, threshold 10, clear threshold 20, for seconds 60, stale after seconds —, 180 | → | state inactive, since —, value 25, held seconds 0 |
| lte: a value at the threshold breaches and waits out the for-duration | samples ×1, op lte, threshold 10, clear threshold —, for seconds 30, stale after seconds —, 20 | → | state pending, since 0, value 10, held seconds 20 |
| no samples at all is no data | , op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 100 | → | state no-data, since —, value —, held seconds 0 |
| samples only after now are ignored, so there is no data yet | samples ×1, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 50 | → | state no-data, since —, value —, held seconds 0 |
| samples after now are ignored | samples ×2, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 30 | → | state firing, since 0, value 95, held seconds 30 |
| a latest sample older than staleAfterSeconds is no data, reporting the stale value | samples ×1, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds 60, 61 | → | state no-data, since —, value 99, held seconds 0 |
| a latest sample exactly staleAfterSeconds old still counts | samples ×1, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds 60, 60 | → | state firing, since 0, value 99, held seconds 60 |
| negative values breach an lt rule | samples ×1, op lt, threshold 0, clear threshold —, for seconds 0, stale after seconds —, 0 | → | state firing, since 0, value -5, held seconds 0 |
| samples out of time order are an error | samples ×2, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 100 | → | error: samples must be in strictly ascending time order |
| a clear level above a gt threshold is an error | samples ×1, op gt, threshold 90, clear threshold 95, for seconds 0, stale after seconds —, 0 | → | error: clearThreshold must be on the non-breaching side of threshold: gt 90, clear 95 |
| a clear level below an lt threshold is an error | samples ×1, op lt, threshold 10, clear threshold 5, for seconds 0, stale after seconds —, 0 | → | error: clearThreshold must be on the non-breaching side of threshold: lt 10, clear 5 |
| a negative for-duration is an error | samples ×1, op gt, threshold 90, clear threshold —, for seconds -1, stale after seconds —, 0 | → | error: forSeconds must not be negative |
| a negative stale limit is an error | samples ×1, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds -1, 0 | → | error: staleAfterSeconds must not be negative |
| an unknown comparison is an error | samples ×1, op eq, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 0 | → | error: unknown comparison: eq |
| a fractional now is an error | samples ×1, op gt, threshold 90, clear threshold —, for seconds 0, stale after seconds —, 1.5 | → | error: now must be a whole number of seconds |
More from the author
1. Only samples with `at <= now` are looked at (a series stored ahead of an evaluation time, as in a replay, does not leak the future). The whole series must still be in strictly ascending time order. 2. The samples are walked in order, carrying a breach: - not breaching: a sample *beyond* `threshold` (by `op`: `gt` is `value > threshold`, `gte` `>=`, `lt` `<`, `lte` `<=`) starts a breach; `since` is that sample's `at`. - breaching: the breach continues while the value is still beyond `clearThreshold` by the same comparison, and clears on the first sample that is not. With `clearThreshold` null it is the threshold itself, so there is no hysteresis. 3. No sample at or before `now`: **no-data**. A `staleAfterSeconds` limit and a latest sample more than that old (`now - at > staleAfterSeconds`): **no-data** too, because a silent exporter is not a healthy one; `value` then reports that stale latest value, and `since` is null. 4. Latest sample breaching: **firing** when `now - since >= forSeconds`, otherwise **pending**. Not breaching: **inactive**.
`heldSeconds` is `now - since` while breaching (pending or firing), else 0.
## Why this shape
The `for` duration follows Prometheus alerting rules: "The optional `for` clause causes Prometheus to wait for a certain duration between first encountering a new expression output vector element and counting an alert as firing for this element", and "Elements that are active, but not firing yet, are in the pending state." Prometheus fires once the alert has been active for at least `for` (`>=`), and with no `for` it fires on the first evaluation; `forSeconds: 0` does the same here.
Hysteresis (a separate clear level) is what stops an alert from flapping when a metric hovers around its threshold: "above 90 to fire, back under 80 to clear". `clearThreshold` must be on the non-breaching side: at or below the threshold for `gt`/`gte`, at or above it for `lt`/`lte`.
A new breach after a clear restarts `since`: the for-duration is about one continuous breach, not the sum of several. Prometheus behaves the same way: an alert whose expression stops returning goes back to inactive and starts over.
Unlike Prometheus this looks at the stored samples, not at a query result per evaluation, so the breach is judged sample by sample; between samples the last value holds.
## Errors
- `samples must be in strictly ascending time order` (from monitor.series-window) - `clearThreshold must be on the non-breaching side of threshold: gt 90, clear 95` - `forSeconds must not be negative` - `staleAfterSeconds must not be negative` - `unknown comparison: eq` - `now must be a whole number of seconds`
## Sources
- Prometheus, "Alerting rules", https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/ (the `for` clause, pending and firing states).
Files
| Path | Bytes |
|---|---|
| README.md | 3,129 |
| impl/python.py | 2,817 |
| impl/rust.rs | 4,383 |
| impl/typescript.ts | 2,721 |
| vectors.json | 7,547 |