Functional Weave
Code in Python

monitor.mttr

MTTR, MTBF, downtime and incident counts for a reporting window, from a list of incidents.

1.0.0 · published 2026-10-03 by charlie · Anterra

Pinned by 18 tests, run in TypeScript, Python and Rust.

What it does

The reliability figures a status report or an SRE review quotes for a window: how many incidents, how much downtime, mean time to recovery (MTTR) and mean time between failures (MTBF). Feed it `monitor.incidents`' output.

## Definitions

For example

  • reliability_stats(incidents ×2, 0, 86,400) → incidents 2, resolved 2, downtime seconds 900, mttr seconds 450, mtbf seconds 42,750, longest seconds 600 two resolved incidents in a day
  • reliability_stats(, 0, 86,400) → incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0 no incidents: no MTTR and no MTBF
  • reliability_stats(incidents ×1, 0, 86,400) → incidents 1, resolved 1, downtime seconds 600, mttr seconds 1,200, mtbf seconds 85,800, longest seconds 1,200 an incident that began before the window: downtime clipped, MTTR uses the full duration

The function

The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.

def reliability_stats(incidents: Sequence[Incident], from_: int, to: int) -> ReliabilityStats
incidentsIncident[]in time order and not overlapping, as monitor.incidents returns them
from_intfirst second of the window, Unix seconds
tointfirst second after the window
returnsReliabilityStats

The type it declares, generated into your project

@dataclass(frozen=True)
class ReliabilityStats:
    """Reliability figures for one window."""

    #: incidents overlapping the window
    incidents: int
    #: of those, how many have ended
    resolved: int
    #: incident time inside the window
    downtime_seconds: int
    #: mean full duration of the resolved incidents, half-up; null when none resolved
    mttr_seconds: Optional[int]
    #: (window - downtime) / incidents, half-up; null when no incidents
    mtbf_seconds: Optional[int]
    #: the longest full duration among them, 0 when none
    longest_seconds: int

Your code names it in one line, in the file that uses it

from fune.monitor.mttr import reliability_stats  # monitor.mttr@^1
impl/python.py · 53 lines · open · raw

Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.

from typing import Optional, Sequence

from .monitor_incidents_types import Incident
from .math_round_div import round_div  ← from math.round-div ^1.0.0 · built alongside by fune
from .monitor_mttr_types import ReliabilityStats


def reliability_stats(incidents: Sequence[Incident], from_: int, to: int) -> ReliabilityStats:
    """Reliability figures for the window from <= t < to. An incident counts
    when it overlaps the window; downtime is clipped to it, but MTTR and the
    longest incident use full durations, because an outage that began last
    month still took as long as it took to fix."""
    for v in (from_, to):
        if not isinstance(v, int) or isinstance(v, bool):
            raise ValueError("from and to must be whole seconds")
    if from_ > to:
        raise ValueError("from must not be after to: %d > %d" % (from_, to))
    prev_end: Optional[int] = None
    count = 0
    resolved = 0
    downtime = 0
    repair_total = 0
    longest = 0
    for x in incidents:
        if x.duration_seconds < 0:
            raise ValueError("durationSeconds must not be negative, received %d" % (x.duration_seconds,))
        if x.end is not None:
            if x.end < x.start:
                raise ValueError("incident end must not be before its start: %d < %d" % (x.end, x.start))
            if x.duration_seconds != x.end - x.start:
                raise ValueError("durationSeconds must equal end - start: %d != %d" % (x.duration_seconds, x.end - x.start))
        # An ongoing incident ends, for now, where its duration says.
        end = x.start + x.duration_seconds
        if prev_end is not None and x.start < prev_end:
            raise ValueError("incidents must be in time order and must not overlap: %d is before %d" % (x.start, prev_end))
        prev_end = end
        if not (x.start < to and (x.start >= from_ or end > from_)):
            continue
        count += 1
        downtime += max(0, min(end, to) - max(x.start, from_))
        if x.end is not None:
            resolved += 1
            repair_total += x.duration_seconds
        if x.duration_seconds > longest:
            longest = x.duration_seconds
    return ReliabilityStats(
        incidents=count,
        resolved=resolved,
        downtime_seconds=downtime,
        mttr_seconds=round_div(repair_total, resolved, "half-up") if resolved > 0 else None,
        mtbf_seconds=round_div(to - from_ - downtime, count, "half-up") if count > 0 else None,
        longest_seconds=longest,
    )

Install

fune build

With that line in your source, in a Python project (language python in fune.project), fune build resolves it and its 2 dependencies, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add monitor.mttr
Download for Python monitor.mttr-1.0.0-python.fune · 13,674 bytes sha256 8c9f593c13a33fea1cc6e66aacbb4e3f21e2c7c7fb19438c5e09f9bbb6a6cbe8

The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./monitor.mttr-1.0.0-python.fune, or fetch it from a terminal with fune pull monitor.mttr@1.0.0:python.

The whole function, every language, is one file too: monitor.mttr-1.0.0.fune, 19,550 bytes, sha256 fe08b30003860518cf3c9e0a257ccace1a5bb8d6a2ae4249d93cc26e353eea9d. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

# fune: before monitor.mttr

after — your function gets the result and the arguments, and returns the final result.

# fune: after monitor.mttr

replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.

# fune: replace math.round-div in monitor.mttr
# fune: replace monitor.incidents in monitor.mttr

step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.mttr --steps.

# fune: step monitor.mttr after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

CaseArgumentsExpected
two resolved incidents in a day incidents ×2, 0, 86,400 → incidents 2, resolved 2, downtime seconds 900, mttr seconds 450, mtbf seconds 42,750, longest seconds 600
no incidents: no MTTR and no MTBF , 0, 86,400 → incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0
an incident that began before the window: downtime clipped, MTTR uses the full duration incidents ×1, 0, 86,400 → incidents 1, resolved 1, downtime seconds 600, mttr seconds 1,200, mtbf seconds 85,800, longest seconds 1,200
an ongoing incident clipped at the window end counts but has no repair time incidents ×1, 0, 86,400 → incidents 1, resolved 0, downtime seconds 400, mttr seconds —, mtbf seconds 86,000, longest seconds 1,000
an incident ending exactly at from and one starting exactly at to are outside the window incidents ×2, 0, 86,400 → incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0
mean repair time and MTBF round half up incidents ×2, 0, 1,000 → incidents 2, resolved 2, downtime seconds 201, mttr seconds 101, mtbf seconds 400, longest seconds 101
a mix of resolved and ongoing: MTTR only averages the resolved one incidents ×2, 0, 1,000 → incidents 2, resolved 1, downtime seconds 400, mttr seconds 100, mtbf seconds 300, longest seconds 300
a zero-length ongoing incident starting at from counts incidents ×1, 0, 100 → incidents 1, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds 100, longest seconds 0
an incident covering the whole window leaves no time between failures incidents ×1, 0, 100 → incidents 1, resolved 1, downtime seconds 100, mttr seconds 300, mtbf seconds 0, longest seconds 300
MTBF of a third rounds down below a half incidents ×3, 0, 100 → incidents 3, resolved 3, downtime seconds 0, mttr seconds 0, mtbf seconds 33, longest seconds 0
Show the other 8 tests
CaseArgumentsExpected
an empty window counts nothing incidents ×1, 50, 50 → incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0
back-to-back incidents touching are allowed incidents ×2, 0, 100 → incidents 2, resolved 2, downtime seconds 30, mttr seconds 15, mtbf seconds 35, longest seconds 20
from after to is an error , 100, 0 → error: from must not be after to
overlapping incidents are an error incidents ×2, 0, 1,000 → error: incidents must be in time order and must not overlap
an incident ending before it starts is an error incidents ×1, 0, 1,000 → error: incident end must not be before its start
a duration that disagrees with end - start is an error incidents ×1, 0, 1,000 → error: durationSeconds must equal end - start
a negative duration is an error incidents ×1, 0, 1,000 → error: durationSeconds must not be negative
a fractional window bound is an error , 0.5, 100 → error: from and to must be whole seconds

More from the author

The window is `from <= t < to`.

- **An incident counts** when it overlaps the window: it starts before `to` and either starts at or after `from` or is still going after `from`. One that ends exactly at `from`, or starts exactly at `to`, belongs to the neighbouring window. An ongoing incident (`end` null) lasts `durationSeconds`, which `monitor.incidents` measured up to its `now`. - **downtimeSeconds** is incident time inside the window only, so back-to-back windows never count the same second twice. - **mttrSeconds** is the mean *full* duration of the resolved incidents that count, rounded half up; null when none has resolved. Full, not clipped: an outage that began before the window still took that long to fix. Ongoing incidents are left out because their repair time is not known yet. This is the usual "total resolution time / number of incidents" definition of mean time to recovery. - **mtbfSeconds** is the operational time in the window divided by the number of incidents, `(to - from - downtimeSeconds) / incidents`, rounded half up; null with no incidents (no failures, no mean between them). This is the reliability-engineering definition: "the sum of the lengths of the operational periods divided by the number of observed failures". Some tools instead measure MTBF from one failure's start to the next's, which includes the repair time; this one does not. - **longestSeconds** is the longest full duration among the counted incidents, 0 when none.

Worked example: incidents of 600 s and 300 s in one day (86400 s) give downtime 900, MTTR 450, MTBF (86400 - 900) / 2 = 42750 s.

## Errors

- `from must not be after to` - `from and to must be whole seconds` - `incidents must be in time order and must not overlap` (touching is fine) - `incident end must not be before its start` - `durationSeconds must equal end - start` for a resolved incident - `durationSeconds must not be negative`

## Sources

- Wikipedia, "Mean time between failures", https://en.wikipedia.org/wiki/Mean_time_between_failures (MTBF = sum of operational periods / number of failures). - Wikipedia, "Mean time to recovery", https://en.wikipedia.org/wiki/Mean_time_to_recovery (MTTR = total resolution time / number of incidents).

Files

PathBytes
README.md2,520
impl/python.py2,444
impl/rust.rs3,330
impl/typescript.ts2,337
vectors.json5,312