monitor.burn-rate-alert
Multiwindow, multi-burn-rate SLO alerting: which rules fire, from event counts per window (Google SRE workbook).
2.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 30 tests, run in TypeScript, Python and Rust.
What it does
SLO alerting the way the Google SRE workbook recommends: multiwindow, multi-burn-rate. Pass the SLO target, the event counts you have for each window length, and the rules (or null for the workbook's), and it says which rules fire and at what severity. An optional minimum number of events keeps a rule quiet on a sample too small to mean anything.
A rule fires when **both** its long window and its short window burn at or above its threshold (see `monitor.burn-rate`). The long window proves enough budget has gone to matter; the short window proves it is still going, so the alert stops soon after the problem does instead of paging for an hour after a five-minute outage.
For example
burnRateAlert(99.9%, windows ×5, —, —)→ firing false, severity —, rules ×3 all quiet: no errors in any window, nothing firesburnRateAlert(99.9%, windows ×5, —, —)→ firing true, severity page, rules ×3 fast burn: 15x over the last hour and 20x in the last 5 minutes pages; the ticket rule fires too, the page winsburnRateAlert(99.9%, windows ×5, —, —)→ firing false, severity —, rules ×3 the outage is over: the hour still burns 15x but the last 5 minutes are clean, so nothing pages
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
export function burnRateAlert(targetBasisPoints: number, windows: readonly WindowCount[], rules: readonly BurnRule[] | null, minEvents: number | null): BurnAlert
| targetBasisPoints | int | the SLO: 9990 = 99.9%; 1 to 9999 |
| windows | WindowCount[] | event counts for each window length the rules use, once each |
| rules | BurnRule[]? | null for the SRE workbook's three rules (Table 5-8) |
| minEvents | int? | a rule fires only when its long window saw at least this many events; null or 0 = no minimum |
| returns | BurnAlert |
The types it declares, generated into your project
/** Events seen over the last windowSeconds. */
export interface WindowCount {
/** at least 1 */
readonly windowSeconds: number;
readonly totalEvents: number;
readonly badEvents: number;
}
/** Fire when both windows burn at or above the threshold. */
export interface BurnRule {
/** e.g. page or ticket */
readonly severity: string;
readonly longWindowSeconds: number;
/** at most longWindowSeconds; the workbook uses 1/12 of it */
readonly shortWindowSeconds: number;
/** threshold, thousandths: 14400 = 14.4x */
readonly burnRateMilli: number;
}
/** One rule, its threshold and both burn rates. */
export interface BurnRuleResult {
readonly severity: string;
readonly longWindowSeconds: number;
readonly shortWindowSeconds: number;
readonly thresholdMilli: number;
/** half-up; firing compares exactly, not this rounded value */
readonly longBurnMilli: number;
readonly shortBurnMilli: number;
/** the long window saw at least minEvents events; true when there is no minimum */
readonly enoughEvents: boolean;
readonly firing: boolean;
}
/** Whether anything fires, the first firing rule's severity, and every rule. */
export interface BurnAlert {
readonly firing: boolean;
/** the first firing rule in rule order; null when none fires */
readonly severity: string | null;
/** in rule order */
readonly rules: readonly BurnRuleResult[];
}
Your code names it in one line, in the file that uses it
import { burnRateAlert } from "#fune/monitor.burn-rate-alert@^2";
Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.
import { burnRate } from "./monitor_burn_rate.ts"; ← from monitor.burn-rate ^1.0.0 · built alongside by fune
import { SRE_WORKBOOK } from "./monitor_burn_rate_alert_data.ts"; ← this capability’s own data, compiled from data/sre-workbook.json into the same file by fune build
import { type BurnAlert, type BurnRule, type BurnRuleResult, type WindowCount } from "./monitor_burn_rate_alert_types.ts";
function whole(value: number): boolean {
return Number.isSafeInteger(value);
}
// bad/total >= threshold/1000 x (10000 - target)/10000, cross-multiplied so a
// burn of 14.3995x never fires a 14.4x rule by rounding up.
function atOrAbove(w: WindowCount, thresholdMilli: number, targetBasisPoints: number): boolean {
if (w.totalEvents === 0) return false;
return BigInt(w.badEvents) * 10000000n >= BigInt(thresholdMilli) * BigInt(w.totalEvents) * BigInt(10000 - targetBasisPoints);
}
/**
* Multiwindow, multi-burn-rate SLO alerting, as the Google SRE workbook
* recommends: a rule fires only when both its long window (enough budget
* spent to matter) and its short window (still happening now) burn at or
* above its threshold. With rules null it uses the workbook's Table 5-8.
*
* minEvents guards small samples: a rule whose long window saw fewer events
* stays quiet, since 2 bad checks out of 11 is a 36x burn of a 99.5% budget
* but proves little. The short window is not guarded: it only confirms the
* burn is still going on.
*/
export function burnRateAlert(
targetBasisPoints: number,
windows: readonly WindowCount[],
rules: readonly BurnRule[] | null,
minEvents: number | null,
): BurnAlert {
if (!whole(targetBasisPoints) || targetBasisPoints < 1 || targetBasisPoints > 9999) {
throw new RangeError(`targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received ${targetBasisPoints}`);
}
if (minEvents !== null && (!whole(minEvents) || minEvents < 0)) {
throw new RangeError(`minEvents must be null or a whole number of at least 0, received ${minEvents}`);
}
const minimum = minEvents ?? 0;
const counts = new Map<number, { count: WindowCount; burn: number }>();
for (const w of windows) {
if (!whole(w.windowSeconds) || w.windowSeconds < 1) throw new RangeError(`windowSeconds must be at least 1, received ${w.windowSeconds}`);
if (counts.has(w.windowSeconds)) throw new RangeError(`duplicate counts for a ${w.windowSeconds}-second window`);
// burnRate checks the counts, so every window is checked, used or not.
counts.set(w.windowSeconds, { count: w, burn: burnRate(targetBasisPoints, w.totalEvents, w.badEvents) });
}
const chosen: readonly BurnRule[] = rules ?? SRE_WORKBOOK.map((r) => ({
severity: r.severity,
longWindowSeconds: r.longWindowSeconds,
shortWindowSeconds: r.shortWindowSeconds,
burnRateMilli: r.burnRateMilli,
}));
if (chosen.length === 0) throw new RangeError("rules must not be empty; pass null for the SRE workbook rules");
const results: BurnRuleResult[] = [];
let severity: string | null = null;
for (const rule of chosen) {
if (rule.severity === "") throw new RangeError("severity must not be empty");
if (!whole(rule.shortWindowSeconds) || rule.shortWindowSeconds < 1) {
throw new RangeError(`shortWindowSeconds must be at least 1, received ${rule.shortWindowSeconds}`);
}
if (!whole(rule.longWindowSeconds) || rule.shortWindowSeconds > rule.longWindowSeconds) {
throw new RangeError(`shortWindowSeconds must not exceed longWindowSeconds: ${rule.shortWindowSeconds} > ${rule.longWindowSeconds}`);
}
if (!whole(rule.burnRateMilli) || rule.burnRateMilli < 1) throw new RangeError(`burnRateMilli must be at least 1, received ${rule.burnRateMilli}`);
const long = counts.get(rule.longWindowSeconds);
if (long === undefined) throw new RangeError(`no counts for a ${rule.longWindowSeconds}-second window`);
const short = counts.get(rule.shortWindowSeconds);
if (short === undefined) throw new RangeError(`no counts for a ${rule.shortWindowSeconds}-second window`);
const enoughEvents = long.count.totalEvents >= minimum;
const firing = enoughEvents && atOrAbove(long.count, rule.burnRateMilli, targetBasisPoints) && atOrAbove(short.count, rule.burnRateMilli, targetBasisPoints);
if (firing && severity === null) severity = rule.severity;
results.push({
severity: rule.severity,
longWindowSeconds: rule.longWindowSeconds,
shortWindowSeconds: rule.shortWindowSeconds,
thresholdMilli: rule.burnRateMilli,
longBurnMilli: long.burn,
shortBurnMilli: short.burn,
enoughEvents,
firing,
});
}
return { firing: severity !== null, severity, rules: results };
}Install
fune build
With that line in your source, in a TypeScript project (language typescript in fune.project), fune build resolves it and its 1 dependency, pins them in fune.lock, downloads only the TypeScript package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:
fune add monitor.burn-rate-alert
The manifest, vectors and README with only the TypeScript implementation. Install it without the registry with fune add ./monitor.burn-rate-alert-2.0.0-typescript.fune, or fetch it from a terminal with fune pull monitor.burn-rate-alert@2.0.0:typescript.
The whole function, every language, is one file too: monitor.burn-rate-alert-2.0.0.fune, 49,170 bytes, sha256 f929270cee119def2ed9395259cd3c8fa3fe53c7ab9b4d545e38f800cb445206. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
// fune: before monitor.burn-rate-alert
after — your function gets the result and the arguments, and returns the final result.
// fune: after monitor.burn-rate-alert
replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.
// fune: replace monitor.burn-rate in monitor.burn-rate-alert
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.burn-rate-alert --steps.
// fune: step monitor.burn-rate-alert after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| all quiet: no errors in any window, nothing fires | 99.9%, windows ×5, —, — | → | firing false, severity —, rules ×3 |
| fast burn: 15x over the last hour and 20x in the last 5 minutes pages; the ticket rule fires too, the page wins | 99.9%, windows ×5, —, — | → | firing true, severity page, rules ×3 |
| the outage is over: the hour still burns 15x but the last 5 minutes are clean, so nothing pages | 99.9%, windows ×5, —, — | → | firing false, severity —, rules ×3 |
| slow burn: 1.2x over three days and 1.5x over six hours raises a ticket, not a page | 99.9%, windows ×5, —, — | → | firing true, severity ticket, rules ×3 |
| 7x for six hours fires the second page rule; windows may come in any order | 99.9%, windows ×5, —, — | → | firing true, severity page, rules ×3 |
| exactly at the threshold fires: 1.44% errors is 14.4x at 99.9% | 99.9%, windows ×2, rules ×1, — | → | firing true, severity page, rules ×1 |
| 14.3995x rounds to 14400 for display but is below 14.4x, so it does not fire | 99.9%, windows ×2, rules ×1, — | → | firing false, severity —, rules ×1 |
| no traffic burns nothing and fires nothing | 99.9%, windows ×2, rules ×1, — | → | firing false, severity —, rules ×1 |
| when several rules fire, the first in rule order gives the severity | 99.9%, windows ×2, rules ×2, — | → | firing true, severity ticket, rules ×2 |
| a 99% SLO allows ten times the errors: 2% errors is only 2x | 99%, windows ×2, rules ×1, — | → | firing false, severity —, rules ×1 |
Show the other 20 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| a single-window rule: short equal to long | 99.9%, windows ×1, rules ×1, — | → | firing true, severity ticket, rules ×1 |
| a new target, 2 bad checks of its first 11 at 99.5%: a 36x burn, but 11 events is under a minimum of 20, so nothing fires | 99.5%, windows ×5, —, 20 | → | firing false, severity —, rules ×3 |
| the same 11 checks with no minimum page, as 1.0.0 did | 99.5%, windows ×5, —, — | → | firing true, severity page, rules ×3 |
| exactly minEvents in the long window is enough; the 5-minute window's 10 is not checked against it | 99.5%, windows ×5, —, 11 | → | firing true, severity page, rules ×3 |
| one event short of the minimum in every long window: nothing fires | 99.5%, windows ×5, —, 12 | → | firing false, severity —, rules ×3 |
| the guard is per rule: the 1-hour rule has 90 events of 100 and stays quiet, the 6-hour rule has 500 and pages | 99.9%, windows ×5, —, 100 | → | firing true, severity page, rules ×3 |
| a minimum of 0 is no minimum, and an empty window still never fires | 99.9%, windows ×1, rules ×1, 0 | → | firing false, severity —, rules ×1 |
| the workbook rules with no counts at all name the first missing window | 99.9%, , —, — | → | error: no counts for a 3600-second window |
| a rule's short window missing from the counts is an error | 99.9%, windows ×1, rules ×1, — | → | error: no counts for a 300-second window |
| two counts for one window length is an error | 99.9%, windows ×3, rules ×1, — | → | error: duplicate counts for a 300-second window |
| a 100% target is an error | 100%, windows ×2, rules ×1, — | → | error: targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received 10000 |
| a zero-length window is an error | 99.9%, windows ×1, rules ×1, — | → | error: windowSeconds must be at least 1, received 0 |
| bad counts are checked even in a window no rule uses | 99.9%, windows ×3, rules ×1, — | → | error: badEvents must not exceed totalEvents: 5 > 4 |
| an empty rule list is an error; null means the workbook rules | 99.9%, windows ×1, , — | → | error: rules must not be empty; pass null for the SRE workbook rules |
| a short window longer than the long window is an error | 99.9%, windows ×2, rules ×1, — | → | error: shortWindowSeconds must not exceed longWindowSeconds: 3600 > 300 |
| a zero threshold is an error | 99.9%, windows ×2, rules ×1, — | → | error: burnRateMilli must be at least 1, received 0 |
| an empty severity is an error | 99.9%, windows ×2, rules ×1, — | → | error: severity must not be empty |
| a negative minimum is an error | 99.9%, windows ×1, —, -1 | → | error: minEvents must be null or a whole number of at least 0, received -1 |
| a fractional minimum is an error | 99.9%, windows ×1, —, 2.5 | → | error: minEvents must be null or a whole number of at least 0, received 2.5 |
| the target is checked before the minimum | 0%, windows ×1, —, -1 | → | error: targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received 0 |
More from the author
## The default rules (rules = null)
Shipped as `data/sre-workbook.json`, from Table 5-8 of the workbook's "Alerting on SLOs" chapter, for a 30-day SLO period:
| Severity | Long window | Short window | Burn rate | Budget consumed | |----------|-------------|--------------|-----------|-----------------| | page | 1 hour | 5 minutes | 14.4 | 2% | | page | 6 hours | 30 minutes | 6 | 5% | | ticket | 3 days | 6 hours | 1 | 10% |
Budget consumed = burn rate x long window / 30 days; the short window is 1/12 of the long one, as the workbook suggests. So the windows to count are 300, 1800, 3600, 21600 and 259200 seconds. The table has no `effective` dates: it is a published recommendation, not a rule that changes in force.
## Small samples: minEvents
Burn rates are ratios, and a ratio of small counts is noise. A new target checked every 30 s has 11 checks after five minutes; if 2 of them failed, its error rate is 18%, which burns a 99.5% budget at 36x and pages under every rule in the table, though nothing much has happened. The workbook names the problem in its section on low-traffic services ("Low-Traffic Services and Error Budget Alerting"): one failure among very few requests can look like a huge burn.
`minEvents` is the usual guard: a rule fires only when its **long** window saw at least that many events. Its `enoughEvents` says whether it did. The short window is not guarded, because it is only there to prove the burn is still happening, and a 5-minute window of a target checked every 30 s never holds more than 10 checks. Since each rule has its own long window, the guard is per rule: with 90 events in the last hour and 500 in the last six, a minimum of 100 silences the 1-hour rule and leaves the 6-hour one free to page.
`null` or `0` means no minimum, which is what 1.0.0 did: every 1.0.0 vector gives the same answer here with `minEvents` null (and `enoughEvents` true).
## Changes from 1.0.0
- A fourth parameter, `minEvents: int?`. - `BurnRuleResult.enoughEvents`.
Both break a caller written for 1.0.0 (one more argument to pass, one more field in the result), hence 2.0.0 rather than 1.1.0: a project on `^1` keeps 1.0.0 until it moves to `^2`. With `minEvents` null the answers are 1.0.0's.
## Decisions
- **Firing is exact.** It compares bad x 10,000,000 against threshold x total x (10000 - target), not the rounded `longBurnMilli`: a burn of 14.3995x shows as `14400` but does not fire a 14.4x rule. - **Severity is the first firing rule in rule order.** List pages before tickets (the workbook table already does), and a fast burn that also trips the ticket rule still pages. - **An empty window (total 0) never fires.** - **Every window is checked**, including ones no rule uses: bad counts in any of them are a bug upstream. - Each rule needs counts for exactly its two window lengths, so the caller counts once per distinct length; a rule with short = long is allowed (a single-window alert). - `minEvents` guards the long window only, and counts every event in it, good or bad. - `rules` = `[]` is an error, since it is far more likely a config mistake than a wish never to alert; null means the workbook rules.
## Errors
- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X` - `minEvents must be null or a whole number of at least 0, received X` - `windowSeconds must be at least 1, received X` - `duplicate counts for a 3600-second window` - the count errors of `monitor.burn-rate` (`badEvents must not exceed totalEvents: B > T`, ...) - `rules must not be empty; pass null for the SRE workbook rules` - `severity must not be empty` - `shortWindowSeconds must be at least 1, received X` - `shortWindowSeconds must not exceed longWindowSeconds: S > L` - `burnRateMilli must be at least 1, received X` - `no counts for a 3600-second window`
## Sources
Google SRE Workbook, chapter 5 "Alerting on SLOs", the multiwindow, multi-burn-rate alerts section and Table 5-8 (https://sre.google/workbook/alerting-on-slos/): page 1h / 5m / 14.4 / 2%; page 6h / 30m / 6 / 5%; ticket 3d / 6h / 1 / 10%; "make the short window 1/12 the duration of the long window"; and the "Low-Traffic Services and Error Budget Alerting" section of the same chapter for the small-sample problem `minEvents` guards against.
Files
| Path | Bytes |
|---|---|
| README.md | 5,115 |
| data/sre-workbook.json | 1,025 |
| impl/python.py | 4,963 |
| impl/rust.rs | 7,231 |
| impl/typescript.ts | 4,577 |
| vectors.json | 18,947 |