monitor.incidents
Turn a series of up, degraded and down checks into incidents, ignoring runs too short to confirm.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 19 tests, run in TypeScript, Python and Rust.
What it does
Finds the incidents in a stored series of checks (`monitor.check-status`'s `Check`): each run of consecutive bad checks is one incident.
## Rules
For example
find_incidents(checks ×5, down, 1, 300)→ ×1 two down checks in a row are one incident ending at the first up checkfind_incidents(checks ×5, down, 2, 300)→ ×1 a run exactly confirmChecks long is an incidentfind_incidents(checks ×5, down, 3, 300)→ a run shorter than confirmChecks is ignored as a flap
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
def find_incidents(checks: Sequence[Check], incident_status: CheckStatus, confirm_checks: int, now: int) -> List[Incident]
| checks | Check[] | in strictly ascending time order |
| incident_status | CheckStatus | down: only down checks are bad; degraded: degraded or down checks are bad |
| confirm_checks | int | how many bad checks in a row make an incident, at least 1 |
| now | int | Unix seconds, not before the last check; ends the duration of an ongoing incident |
| returns | Incident[] |
The type it declares, generated into your project
@dataclass(frozen=True)
class Incident:
"""One run of bad checks, from the first bad check to the first good one after it."""
#: at of the first bad check
start: int
#: at of the first good check after the run; null while ongoing
end: Optional[int]
#: (end, or now while ongoing) - start
duration_seconds: int
#: down if any check in the run was down, else degraded
worst: CheckStatus
#: how many bad checks the run holds
checks: int
Your code names it in one line, in the file that uses it
from fune.monitor.incidents import find_incidents # monitor.incidents@^1
from typing import List, Optional, Sequence
from .monitor_check_status_types import Check, CheckStatus
from .monitor_incidents_types import Incident
def find_incidents(checks: Sequence[Check], incident_status: CheckStatus, confirm_checks: int, now: int) -> List[Incident]:
"""The incidents in a check series. A run of consecutive bad checks is
one incident, from its first bad check to the first good check after it,
but only when it holds at least confirm_checks checks: a single failed
probe is noise, not an outage, and would otherwise page and skew MTTR."""
if not isinstance(now, int) or isinstance(now, bool):
raise ValueError("now must be a whole number of seconds")
if incident_status not in ("down", "degraded"):
raise ValueError("incidentStatus must be down or degraded, received %s" % (incident_status,))
if not isinstance(confirm_checks, int) or isinstance(confirm_checks, bool) or confirm_checks < 1:
raise ValueError("confirmChecks must be at least 1, received %r" % (confirm_checks,))
for i, c in enumerate(checks):
if c.status not in ("up", "degraded", "down"):
raise ValueError("unknown check status: %s" % (c.status,))
if i > 0 and c.at <= checks[i - 1].at:
raise ValueError("checks must be in strictly ascending time order: %d follows %d" % (c.at, checks[i - 1].at))
if len(checks) > 0 and now < checks[-1].at:
raise ValueError("now %d is before the last check at %d" % (now, checks[-1].at))
def bad(s: str) -> bool:
return s == "down" or (incident_status == "degraded" and s == "degraded")
out: List[Incident] = []
i = 0
n = len(checks)
while i < n:
if not bad(checks[i].status):
i += 1
continue
start = checks[i].at
count = 0
worst = "degraded"
while i < n and bad(checks[i].status):
if checks[i].status == "down":
worst = "down"
count += 1
i += 1
end: Optional[int] = checks[i].at if i < n else None
if count >= confirm_checks:
out.append(Incident(start=start, end=end, duration_seconds=(now if end is None else end) - start, worst=worst, checks=count))
return outInstall
fune build
With that line in your source, in a Python project (language python in fune.project), fune build resolves it and its 1 dependency, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:
fune add monitor.incidents
The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./monitor.incidents-1.0.0-python.fune, or fetch it from a terminal with fune pull monitor.incidents@1.0.0:python.
The whole function, every language, is one file too: monitor.incidents-1.0.0.fune, 19,417 bytes, sha256 4664c6be7bf3e0139623ab73f3d74becc8b401bdee0de17e21fa7824a05124a8. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
# fune: before monitor.incidents
after — your function gets the result and the arguments, and returns the final result.
# fune: after monitor.incidents
replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.
# fune: replace monitor.check-status in monitor.incidents
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.incidents --steps.
# fune: step monitor.incidents after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| two down checks in a row are one incident ending at the first up check | checks ×5, down, 1, 300 | → | ×1 |
| a run exactly confirmChecks long is an incident | checks ×5, down, 2, 300 | → | ×1 |
| a run shorter than confirmChecks is ignored as a flap | checks ×5, down, 3, 300 | → | |
| a run still bad at the last check is ongoing, measured to now | checks ×3, down, 1, 200 | → | ×1 |
| degraded mode joins degraded and down checks into one incident, worst down | checks ×5, degraded, 1, 240 | → | ×1 |
| down mode: a degraded check is not bad, so it ends the down incident | checks ×5, down, 1, 240 | → | ×1 |
| a degraded-only incident has worst degraded | checks ×3, degraded, 1, 120 | → | ×1 |
| short blips around a real outage are dropped, and an unconfirmed ongoing run too | checks ×8, down, 2, 480 | → | ×1 |
| no checks, no incidents | , down, 1, 1,000 | → | |
| a series that starts bad starts its incident at the first check | checks ×2, down, 1, 60 | → | ×1 |
Show the other 9 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| all up has no incidents | checks ×3, down, 1, 180 | → | |
| an ongoing incident with now at its only check lasts 0 seconds | checks ×1, down, 1, 100 | → | ×1 |
| two separate incidents come out in time order | checks ×4, down, 1, 45 | → | ×2 |
| checks out of time order are an error | checks ×2, down, 1, 100 | → | error: checks must be in strictly ascending time order |
| an unknown check status is an error | checks ×1, down, 1, 100 | → | error: unknown check status: offline |
| up is not an incident status | checks ×5, up, 1, 300 | → | error: incidentStatus must be down or degraded, received up |
| confirmChecks 0 is an error | checks ×5, down, 0, 300 | → | error: confirmChecks must be at least 1 |
| now before the last check is an error | checks ×5, down, 1, 200 | → | error: now 200 is before the last check at 240 |
| a fractional now is an error | checks ×5, down, 1, 300.5 | → | error: now must be a whole number of seconds |
More from the author
- **Bad** depends on `incidentStatus`: with `down`, only `down` checks are bad (a slow site is not an outage); with `degraded`, `degraded` and `down` both are, and one incident can mix them. `up` is refused: an "incident of being up" is a configuration mistake. - A run counts only when it has at least `confirmChecks` checks. This is flap suppression, the same idea as "alert after N consecutive failures" in most uptime monitors: one lost probe does not make an outage, and would drag MTTR down if it did. Shorter runs are dropped entirely, including an unconfirmed run still in progress at the last check. - `start` is the `at` of the first bad check; `end` is the `at` of the first check after the run that is not bad (with `incidentStatus: down`, a `degraded` check ends a down incident). An incident still bad at the last check is ongoing: `end` is null and `durationSeconds` runs to `now`. - The start is the first check that *saw* the failure, not a guess at when it really began between two checks; and the end is the first check that saw recovery. So durations are accurate to the check interval, and err long. - `worst` is `down` if any check in the run was down, else `degraded`. - `checks` counts the bad checks in the run.
Incidents come out in time order and never overlap, which is what `monitor.mttr` expects.
## Errors
- `checks must be in strictly ascending time order` - `unknown check status: X` - `incidentStatus must be down or degraded, received up` - `confirmChecks must be at least 1` - `now 200 is before the last check at 240`: now ends ongoing incidents, so it cannot be earlier than the data. - `now must be a whole number of seconds`
Files
| Path | Bytes |
|---|---|
| README.md | 1,863 |
| impl/python.py | 2,276 |
| impl/rust.rs | 3,891 |
| impl/typescript.ts | 2,253 |
| vectors.json | 5,284 |