monitor.mttr
MTTR, MTBF, downtime and incident counts for a reporting window, from a list of incidents.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 18 tests, run in TypeScript, Python and Rust.
What it does
The reliability figures a status report or an SRE review quotes for a window: how many incidents, how much downtime, mean time to recovery (MTTR) and mean time between failures (MTBF). Feed it `monitor.incidents`' output.
## Definitions
For example
reliability_stats(incidents ×2, 0, 86,400)→ incidents 2, resolved 2, downtime seconds 900, mttr seconds 450, mtbf seconds 42,750, longest seconds 600 two resolved incidents in a dayreliability_stats(, 0, 86,400)→ incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0 no incidents: no MTTR and no MTBFreliability_stats(incidents ×1, 0, 86,400)→ incidents 1, resolved 1, downtime seconds 600, mttr seconds 1,200, mtbf seconds 85,800, longest seconds 1,200 an incident that began before the window: downtime clipped, MTTR uses the full duration
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
pub fn reliability_stats(incidents: &[Incident], from: i64, to: i64) -> ReliabilityStats
| incidents | Incident[] | in time order and not overlapping, as monitor.incidents returns them |
| from | int | first second of the window, Unix seconds |
| to | int | first second after the window |
| returns | ReliabilityStats |
The type it declares, generated into your project
/// Reliability figures for one window.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct ReliabilityStats {
/// incidents overlapping the window
pub incidents: i64,
/// of those, how many have ended
pub resolved: i64,
/// incident time inside the window
pub downtime_seconds: i64,
/// mean full duration of the resolved incidents, half-up; null when none resolved
pub mttr_seconds: Option<i64>,
/// (window - downtime) / incidents, half-up; null when no incidents
pub mtbf_seconds: Option<i64>,
/// the longest full duration among them, 0 when none
pub longest_seconds: i64,
}
Your code names it in one line, in the file that uses it
fune!(monitor.mttr@^1); // then call reliability_stats(…)
Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.
use super::funejson::Value; ← the fune runtime: the JSON value the test vectors use; fune build keeps it only where a signature takes one
use super::math_round_div::round_div; ← from math.round-div ^1.0.0 · built alongside by fune
use super::monitor_incidents::{incidents_from_value, Incident}; ← from monitor.incidents ^1.0.0 · built alongside by fune
/// Reliability figures for the window from <= t < to. An incident counts when
/// it overlaps the window; downtime is clipped to it, but MTTR and the
/// longest incident use full durations, because an outage that began last
/// month still took as long as it took to fix.
///
/// # Panics
/// Panics when from is after to, on an inconsistent incident, or on
/// incidents out of order or overlapping.
pub fn reliability_stats(incidents: &[Incident], from: i64, to: i64) -> ReliabilityStats {
if from > to {
panic!("from must not be after to: {} > {}", from, to);
}
let mut prev_end: Option<i64> = None;
let (mut count, mut resolved, mut downtime, mut repair_total, mut longest) = (0i64, 0i64, 0i64, 0i64, 0i64);
for x in incidents {
if x.duration_seconds < 0 {
panic!("durationSeconds must not be negative, received {}", x.duration_seconds);
}
if let Some(e) = x.end {
if e < x.start {
panic!("incident end must not be before its start: {} < {}", e, x.start);
}
if x.duration_seconds != e - x.start {
panic!("durationSeconds must equal end - start: {} != {}", x.duration_seconds, e - x.start);
}
}
// An ongoing incident ends, for now, where its duration says.
let end = x.start + x.duration_seconds;
if let Some(p) = prev_end {
if x.start < p {
panic!("incidents must be in time order and must not overlap: {} is before {}", x.start, p);
}
}
prev_end = Some(end);
if !(x.start < to && (x.start >= from || end > from)) {
continue;
}
count += 1;
downtime += (end.min(to) - x.start.max(from)).max(0);
if x.end.is_some() {
resolved += 1;
repair_total += x.duration_seconds;
}
longest = longest.max(x.duration_seconds);
}
ReliabilityStats {
incidents: count,
resolved,
downtime_seconds: downtime,
mttr_seconds: if resolved > 0 { Some(round_div(repair_total, resolved, "half-up")) } else { None },
mtbf_seconds: if count > 0 { Some(round_div(to - from - downtime, count, "half-up")) } else { None },
longest_seconds: longest,
}
}
fn opt_to_value(v: Option<i64>) -> Value {
match v {
Some(i) => Value::Int(i),
None => Value::Null,
}
}
pub fn reliability_stats_to_value(r: &ReliabilityStats) -> Value {
Value::obj(vec![
("incidents", Value::Int(r.incidents)),
("resolved", Value::Int(r.resolved)),
("downtimeSeconds", Value::Int(r.downtime_seconds)),
("mttrSeconds", opt_to_value(r.mttr_seconds)),
("mtbfSeconds", opt_to_value(r.mtbf_seconds)),
("longestSeconds", Value::Int(r.longest_seconds)),
])
}
fn whole(v: &Value) -> i64 {
match v {
Value::Int(i) => *i,
_ => panic!("from and to must be whole seconds"),
}
}
pub fn fune_vector(args: &[Value]) -> Value {
let incidents = incidents_from_value(&args[0]);
reliability_stats_to_value(&reliability_stats(&incidents, whole(&args[1]), whole(&args[2])))
}Install
fune build
With that line in your source, in a Rust project (language rust in fune.project), fune build resolves it and its 2 dependencies, pins them in fune.lock, downloads only the Rust package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. A crate’s build.rs runs it before every compile. Or pin a range in fune.project and build in one step:
fune add monitor.mttr
The manifest, vectors and README with only the Rust implementation. Install it without the registry with fune add ./monitor.mttr-1.0.0-rust.fune, or fetch it from a terminal with fune pull monitor.mttr@1.0.0:rust.
The whole function, every language, is one file too: monitor.mttr-1.0.0.fune, 19,550 bytes, sha256 fe08b30003860518cf3c9e0a257ccace1a5bb8d6a2ae4249d93cc26e353eea9d. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
// fune: before monitor.mttr
after — your function gets the result and the arguments, and returns the final result.
// fune: after monitor.mttr
replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.
// fune: replace math.round-div in monitor.mttr
// fune: replace monitor.incidents in monitor.mttr
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.mttr --steps.
// fune: step monitor.mttr after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| two resolved incidents in a day | incidents ×2, 0, 86,400 | → | incidents 2, resolved 2, downtime seconds 900, mttr seconds 450, mtbf seconds 42,750, longest seconds 600 |
| no incidents: no MTTR and no MTBF | , 0, 86,400 | → | incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0 |
| an incident that began before the window: downtime clipped, MTTR uses the full duration | incidents ×1, 0, 86,400 | → | incidents 1, resolved 1, downtime seconds 600, mttr seconds 1,200, mtbf seconds 85,800, longest seconds 1,200 |
| an ongoing incident clipped at the window end counts but has no repair time | incidents ×1, 0, 86,400 | → | incidents 1, resolved 0, downtime seconds 400, mttr seconds —, mtbf seconds 86,000, longest seconds 1,000 |
| an incident ending exactly at from and one starting exactly at to are outside the window | incidents ×2, 0, 86,400 | → | incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0 |
| mean repair time and MTBF round half up | incidents ×2, 0, 1,000 | → | incidents 2, resolved 2, downtime seconds 201, mttr seconds 101, mtbf seconds 400, longest seconds 101 |
| a mix of resolved and ongoing: MTTR only averages the resolved one | incidents ×2, 0, 1,000 | → | incidents 2, resolved 1, downtime seconds 400, mttr seconds 100, mtbf seconds 300, longest seconds 300 |
| a zero-length ongoing incident starting at from counts | incidents ×1, 0, 100 | → | incidents 1, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds 100, longest seconds 0 |
| an incident covering the whole window leaves no time between failures | incidents ×1, 0, 100 | → | incidents 1, resolved 1, downtime seconds 100, mttr seconds 300, mtbf seconds 0, longest seconds 300 |
| MTBF of a third rounds down below a half | incidents ×3, 0, 100 | → | incidents 3, resolved 3, downtime seconds 0, mttr seconds 0, mtbf seconds 33, longest seconds 0 |
Show the other 8 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| an empty window counts nothing | incidents ×1, 50, 50 | → | incidents 0, resolved 0, downtime seconds 0, mttr seconds —, mtbf seconds —, longest seconds 0 |
| back-to-back incidents touching are allowed | incidents ×2, 0, 100 | → | incidents 2, resolved 2, downtime seconds 30, mttr seconds 15, mtbf seconds 35, longest seconds 20 |
| from after to is an error | , 100, 0 | → | error: from must not be after to |
| overlapping incidents are an error | incidents ×2, 0, 1,000 | → | error: incidents must be in time order and must not overlap |
| an incident ending before it starts is an error | incidents ×1, 0, 1,000 | → | error: incident end must not be before its start |
| a duration that disagrees with end - start is an error | incidents ×1, 0, 1,000 | → | error: durationSeconds must equal end - start |
| a negative duration is an error | incidents ×1, 0, 1,000 | → | error: durationSeconds must not be negative |
| a fractional window bound is an error | , 0.5, 100 | → | error: from and to must be whole seconds |
More from the author
The window is `from <= t < to`.
- **An incident counts** when it overlaps the window: it starts before `to` and either starts at or after `from` or is still going after `from`. One that ends exactly at `from`, or starts exactly at `to`, belongs to the neighbouring window. An ongoing incident (`end` null) lasts `durationSeconds`, which `monitor.incidents` measured up to its `now`. - **downtimeSeconds** is incident time inside the window only, so back-to-back windows never count the same second twice. - **mttrSeconds** is the mean *full* duration of the resolved incidents that count, rounded half up; null when none has resolved. Full, not clipped: an outage that began before the window still took that long to fix. Ongoing incidents are left out because their repair time is not known yet. This is the usual "total resolution time / number of incidents" definition of mean time to recovery. - **mtbfSeconds** is the operational time in the window divided by the number of incidents, `(to - from - downtimeSeconds) / incidents`, rounded half up; null with no incidents (no failures, no mean between them). This is the reliability-engineering definition: "the sum of the lengths of the operational periods divided by the number of observed failures". Some tools instead measure MTBF from one failure's start to the next's, which includes the repair time; this one does not. - **longestSeconds** is the longest full duration among the counted incidents, 0 when none.
Worked example: incidents of 600 s and 300 s in one day (86400 s) give downtime 900, MTTR 450, MTBF (86400 - 900) / 2 = 42750 s.
## Errors
- `from must not be after to` - `from and to must be whole seconds` - `incidents must be in time order and must not overlap` (touching is fine) - `incident end must not be before its start` - `durationSeconds must equal end - start` for a resolved incident - `durationSeconds must not be negative`
## Sources
- Wikipedia, "Mean time between failures", https://en.wikipedia.org/wiki/Mean_time_between_failures (MTBF = sum of operational periods / number of failures). - Wikipedia, "Mean time to recovery", https://en.wikipedia.org/wiki/Mean_time_to_recovery (MTTR = total resolution time / number of incidents).
Files
| Path | Bytes |
|---|---|
| README.md | 2,520 |
| impl/python.py | 2,444 |
| impl/rust.rs | 3,330 |
| impl/typescript.ts | 2,337 |
| vectors.json | 5,312 |