Functional Weave
Code in Python

monitor.burn-rate-alert

Multiwindow, multi-burn-rate SLO alerting: which rules fire, from event counts per window (Google SRE workbook).

2.0.0 · published 2026-10-03 by charlie · Anterra

Pinned by 30 tests, run in TypeScript, Python and Rust.

What it does

SLO alerting the way the Google SRE workbook recommends: multiwindow, multi-burn-rate. Pass the SLO target, the event counts you have for each window length, and the rules (or null for the workbook's), and it says which rules fire and at what severity. An optional minimum number of events keeps a rule quiet on a sample too small to mean anything.

A rule fires when **both** its long window and its short window burn at or above its threshold (see `monitor.burn-rate`). The long window proves enough budget has gone to matter; the short window proves it is still going, so the alert stops soon after the problem does instead of paging for an hour after a five-minute outage.

For example

  • burn_rate_alert(99.9%, windows ×5, —, —) → firing false, severity —, rules ×3 all quiet: no errors in any window, nothing fires
  • burn_rate_alert(99.9%, windows ×5, —, —) → firing true, severity page, rules ×3 fast burn: 15x over the last hour and 20x in the last 5 minutes pages; the ticket rule fires too, the page wins
  • burn_rate_alert(99.9%, windows ×5, —, —) → firing false, severity —, rules ×3 the outage is over: the hour still burns 15x but the last 5 minutes are clean, so nothing pages

The function

The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.

def burn_rate_alert(target_basis_points: int, windows: Sequence[WindowCount], rules: Optional[Sequence[BurnRule]], min_events: Optional[int]) -> BurnAlert
target_basis_pointsintthe SLO: 9990 = 99.9%; 1 to 9999
windowsWindowCount[]event counts for each window length the rules use, once each
rulesBurnRule[]?null for the SRE workbook's three rules (Table 5-8)
min_eventsint?a rule fires only when its long window saw at least this many events; null or 0 = no minimum
returnsBurnAlert

The types it declares, generated into your project

@dataclass(frozen=True)
class WindowCount:
    """Events seen over the last windowSeconds."""

    #: at least 1
    window_seconds: int
    total_events: int
    bad_events: int

@dataclass(frozen=True)
class BurnRule:
    """Fire when both windows burn at or above the threshold."""

    #: e.g. page or ticket
    severity: str
    long_window_seconds: int
    #: at most longWindowSeconds; the workbook uses 1/12 of it
    short_window_seconds: int
    #: threshold, thousandths: 14400 = 14.4x
    burn_rate_milli: int

@dataclass(frozen=True)
class BurnRuleResult:
    """One rule, its threshold and both burn rates."""

    severity: str
    long_window_seconds: int
    short_window_seconds: int
    threshold_milli: int
    #: half-up; firing compares exactly, not this rounded value
    long_burn_milli: int
    short_burn_milli: int
    #: the long window saw at least minEvents events; true when there is no minimum
    enough_events: bool
    firing: bool

@dataclass(frozen=True)
class BurnAlert:
    """Whether anything fires, the first firing rule's severity, and every rule."""

    firing: bool
    #: the first firing rule in rule order; null when none fires
    severity: Optional[str]
    #: in rule order
    rules: List[BurnRuleResult]

Your code names it in one line, in the file that uses it

from fune.monitor.burn_rate_alert import burn_rate_alert  # monitor.burn-rate-alert@^2
impl/python.py · 85 lines · open · raw

Imports name this capability’s declared dependencies, which fune builds next to it in your project; each one links to its page.

from typing import Dict, List, Optional, Sequence, Tuple

from .monitor_burn_rate import burn_rate  ← from monitor.burn-rate ^1.0.0 · built alongside by fune
from .monitor_burn_rate_alert_data import SRE_WORKBOOK  ← this capability’s own data, compiled from data/sre-workbook.json into the same file by fune build
from .monitor_burn_rate_alert_types import BurnAlert, BurnRule, BurnRuleResult, WindowCount


def _whole(value: object) -> bool:
    return isinstance(value, int) and not isinstance(value, bool)


def _at_or_above(w: WindowCount, threshold_milli: int, target_basis_points: int) -> bool:
    # bad/total >= threshold/1000 x (10000 - target)/10000, cross-multiplied so a
    # burn of 14.3995x never fires a 14.4x rule by rounding up.
    if w.total_events == 0:
        return False
    return w.bad_events * 10000000 >= threshold_milli * w.total_events * (10000 - target_basis_points)


def burn_rate_alert(target_basis_points: int, windows: Sequence[WindowCount], rules: Optional[Sequence[BurnRule]], min_events: Optional[int]) -> BurnAlert:
    """Multiwindow, multi-burn-rate SLO alerting, as the Google SRE workbook
    recommends: a rule fires only when both its long window (enough budget
    spent to matter) and its short window (still happening now) burn at or
    above its threshold. With rules None it uses the workbook's Table 5-8.

    min_events guards small samples: a rule whose long window saw fewer
    events stays quiet, since 2 bad checks out of 11 is a 36x burn of a
    99.5% budget but proves little. The short window is not guarded: it
    only confirms the burn is still going on."""
    if not _whole(target_basis_points) or target_basis_points < 1 or target_basis_points > 9999:
        raise ValueError("targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received %s" % (target_basis_points,))
    if min_events is not None and (not _whole(min_events) or min_events < 0):
        raise ValueError("minEvents must be null or a whole number of at least 0, received %s" % (min_events,))
    minimum = 0 if min_events is None else min_events
    counts: Dict[int, Tuple[WindowCount, int]] = {}
    for w in windows:
        if not _whole(w.window_seconds) or w.window_seconds < 1:
            raise ValueError("windowSeconds must be at least 1, received %s" % (w.window_seconds,))
        if w.window_seconds in counts:
            raise ValueError("duplicate counts for a %d-second window" % w.window_seconds)
        # burn_rate checks the counts, so every window is checked, used or not.
        counts[w.window_seconds] = (w, burn_rate(target_basis_points, w.total_events, w.bad_events))
    chosen: Sequence[BurnRule] = rules if rules is not None else [
        BurnRule(
            severity=r.severity,
            long_window_seconds=r.long_window_seconds,
            short_window_seconds=r.short_window_seconds,
            burn_rate_milli=r.burn_rate_milli,
        )
        for r in SRE_WORKBOOK
    ]
    if len(chosen) == 0:
        raise ValueError("rules must not be empty; pass null for the SRE workbook rules")
    results: List[BurnRuleResult] = []
    severity: Optional[str] = None
    for rule in chosen:
        if rule.severity == "":
            raise ValueError("severity must not be empty")
        if not _whole(rule.short_window_seconds) or rule.short_window_seconds < 1:
            raise ValueError("shortWindowSeconds must be at least 1, received %s" % (rule.short_window_seconds,))
        if not _whole(rule.long_window_seconds) or rule.short_window_seconds > rule.long_window_seconds:
            raise ValueError("shortWindowSeconds must not exceed longWindowSeconds: %s > %s" % (rule.short_window_seconds, rule.long_window_seconds))
        if not _whole(rule.burn_rate_milli) or rule.burn_rate_milli < 1:
            raise ValueError("burnRateMilli must be at least 1, received %s" % (rule.burn_rate_milli,))
        if rule.long_window_seconds not in counts:
            raise ValueError("no counts for a %d-second window" % rule.long_window_seconds)
        if rule.short_window_seconds not in counts:
            raise ValueError("no counts for a %d-second window" % rule.short_window_seconds)
        long_count, long_burn = counts[rule.long_window_seconds]
        short_count, short_burn = counts[rule.short_window_seconds]
        enough = long_count.total_events >= minimum
        firing = enough and _at_or_above(long_count, rule.burn_rate_milli, target_basis_points) and _at_or_above(short_count, rule.burn_rate_milli, target_basis_points)
        if firing and severity is None:
            severity = rule.severity
        results.append(BurnRuleResult(
            severity=rule.severity,
            long_window_seconds=rule.long_window_seconds,
            short_window_seconds=rule.short_window_seconds,
            threshold_milli=rule.burn_rate_milli,
            long_burn_milli=long_burn,
            short_burn_milli=short_burn,
            enough_events=enough,
            firing=firing,
        ))
    return BurnAlert(firing=severity is not None, severity=severity, rules=results)

Install

fune build

With that line in your source, in a Python project (language python in fune.project), fune build resolves it and its 1 dependency, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:

fune add monitor.burn-rate-alert
Download for Python monitor.burn-rate-alert-2.0.0-python.fune · 37,007 bytes sha256 aeb140c3d47f35ab42ce321fa11aeb3c2b4c0b36e8f64ecb2911476553bb85d7

The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./monitor.burn-rate-alert-2.0.0-python.fune, or fetch it from a terminal with fune pull monitor.burn-rate-alert@2.0.0:python.

The whole function, every language, is one file too: monitor.burn-rate-alert-2.0.0.fune, 49,170 bytes, sha256 f929270cee119def2ed9395259cd3c8fa3fe53c7ab9b4d545e38f800cb445206. It installs into a project of any language.

Customise it in your app

The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.

before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.

# fune: before monitor.burn-rate-alert

after — your function gets the result and the arguments, and returns the final result.

# fune: after monitor.burn-rate-alert

replace — inside this capability’s code only, calls to a dependency go to your function, with the same signature. Other capabilities that use it are unaffected; write in * to replace it everywhere.

# fune: replace monitor.burn-rate in monitor.burn-rate-alert

step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.burn-rate-alert --steps.

# fune: step monitor.burn-rate-alert after <n|label>

Tests

A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.

CaseArgumentsExpected
all quiet: no errors in any window, nothing fires 99.9%, windows ×5, —, — → firing false, severity —, rules ×3
fast burn: 15x over the last hour and 20x in the last 5 minutes pages; the ticket rule fires too, the page wins 99.9%, windows ×5, —, — → firing true, severity page, rules ×3
the outage is over: the hour still burns 15x but the last 5 minutes are clean, so nothing pages 99.9%, windows ×5, —, — → firing false, severity —, rules ×3
slow burn: 1.2x over three days and 1.5x over six hours raises a ticket, not a page 99.9%, windows ×5, —, — → firing true, severity ticket, rules ×3
7x for six hours fires the second page rule; windows may come in any order 99.9%, windows ×5, —, — → firing true, severity page, rules ×3
exactly at the threshold fires: 1.44% errors is 14.4x at 99.9% 99.9%, windows ×2, rules ×1, — → firing true, severity page, rules ×1
14.3995x rounds to 14400 for display but is below 14.4x, so it does not fire 99.9%, windows ×2, rules ×1, — → firing false, severity —, rules ×1
no traffic burns nothing and fires nothing 99.9%, windows ×2, rules ×1, — → firing false, severity —, rules ×1
when several rules fire, the first in rule order gives the severity 99.9%, windows ×2, rules ×2, — → firing true, severity ticket, rules ×2
a 99% SLO allows ten times the errors: 2% errors is only 2x 99%, windows ×2, rules ×1, — → firing false, severity —, rules ×1
Show the other 20 tests
CaseArgumentsExpected
a single-window rule: short equal to long 99.9%, windows ×1, rules ×1, — → firing true, severity ticket, rules ×1
a new target, 2 bad checks of its first 11 at 99.5%: a 36x burn, but 11 events is under a minimum of 20, so nothing fires 99.5%, windows ×5, —, 20 → firing false, severity —, rules ×3
the same 11 checks with no minimum page, as 1.0.0 did 99.5%, windows ×5, —, — → firing true, severity page, rules ×3
exactly minEvents in the long window is enough; the 5-minute window's 10 is not checked against it 99.5%, windows ×5, —, 11 → firing true, severity page, rules ×3
one event short of the minimum in every long window: nothing fires 99.5%, windows ×5, —, 12 → firing false, severity —, rules ×3
the guard is per rule: the 1-hour rule has 90 events of 100 and stays quiet, the 6-hour rule has 500 and pages 99.9%, windows ×5, —, 100 → firing true, severity page, rules ×3
a minimum of 0 is no minimum, and an empty window still never fires 99.9%, windows ×1, rules ×1, 0 → firing false, severity —, rules ×1
the workbook rules with no counts at all name the first missing window 99.9%, , —, — → error: no counts for a 3600-second window
a rule's short window missing from the counts is an error 99.9%, windows ×1, rules ×1, — → error: no counts for a 300-second window
two counts for one window length is an error 99.9%, windows ×3, rules ×1, — → error: duplicate counts for a 300-second window
a 100% target is an error 100%, windows ×2, rules ×1, — → error: targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received 10000
a zero-length window is an error 99.9%, windows ×1, rules ×1, — → error: windowSeconds must be at least 1, received 0
bad counts are checked even in a window no rule uses 99.9%, windows ×3, rules ×1, — → error: badEvents must not exceed totalEvents: 5 > 4
an empty rule list is an error; null means the workbook rules 99.9%, windows ×1, , — → error: rules must not be empty; pass null for the SRE workbook rules
a short window longer than the long window is an error 99.9%, windows ×2, rules ×1, — → error: shortWindowSeconds must not exceed longWindowSeconds: 3600 > 300
a zero threshold is an error 99.9%, windows ×2, rules ×1, — → error: burnRateMilli must be at least 1, received 0
an empty severity is an error 99.9%, windows ×2, rules ×1, — → error: severity must not be empty
a negative minimum is an error 99.9%, windows ×1, —, -1 → error: minEvents must be null or a whole number of at least 0, received -1
a fractional minimum is an error 99.9%, windows ×1, —, 2.5 → error: minEvents must be null or a whole number of at least 0, received 2.5
the target is checked before the minimum 0%, windows ×1, —, -1 → error: targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received 0

More from the author

## The default rules (rules = null)

Shipped as `data/sre-workbook.json`, from Table 5-8 of the workbook's "Alerting on SLOs" chapter, for a 30-day SLO period:

| Severity | Long window | Short window | Burn rate | Budget consumed | |----------|-------------|--------------|-----------|-----------------| | page | 1 hour | 5 minutes | 14.4 | 2% | | page | 6 hours | 30 minutes | 6 | 5% | | ticket | 3 days | 6 hours | 1 | 10% |

Budget consumed = burn rate x long window / 30 days; the short window is 1/12 of the long one, as the workbook suggests. So the windows to count are 300, 1800, 3600, 21600 and 259200 seconds. The table has no `effective` dates: it is a published recommendation, not a rule that changes in force.

## Small samples: minEvents

Burn rates are ratios, and a ratio of small counts is noise. A new target checked every 30 s has 11 checks after five minutes; if 2 of them failed, its error rate is 18%, which burns a 99.5% budget at 36x and pages under every rule in the table, though nothing much has happened. The workbook names the problem in its section on low-traffic services ("Low-Traffic Services and Error Budget Alerting"): one failure among very few requests can look like a huge burn.

`minEvents` is the usual guard: a rule fires only when its **long** window saw at least that many events. Its `enoughEvents` says whether it did. The short window is not guarded, because it is only there to prove the burn is still happening, and a 5-minute window of a target checked every 30 s never holds more than 10 checks. Since each rule has its own long window, the guard is per rule: with 90 events in the last hour and 500 in the last six, a minimum of 100 silences the 1-hour rule and leaves the 6-hour one free to page.

`null` or `0` means no minimum, which is what 1.0.0 did: every 1.0.0 vector gives the same answer here with `minEvents` null (and `enoughEvents` true).

## Changes from 1.0.0

- A fourth parameter, `minEvents: int?`. - `BurnRuleResult.enoughEvents`.

Both break a caller written for 1.0.0 (one more argument to pass, one more field in the result), hence 2.0.0 rather than 1.1.0: a project on `^1` keeps 1.0.0 until it moves to `^2`. With `minEvents` null the answers are 1.0.0's.

## Decisions

- **Firing is exact.** It compares bad x 10,000,000 against threshold x total x (10000 - target), not the rounded `longBurnMilli`: a burn of 14.3995x shows as `14400` but does not fire a 14.4x rule. - **Severity is the first firing rule in rule order.** List pages before tickets (the workbook table already does), and a fast burn that also trips the ticket rule still pages. - **An empty window (total 0) never fires.** - **Every window is checked**, including ones no rule uses: bad counts in any of them are a bug upstream. - Each rule needs counts for exactly its two window lengths, so the caller counts once per distinct length; a rule with short = long is allowed (a single-window alert). - `minEvents` guards the long window only, and counts every event in it, good or bad. - `rules` = `[]` is an error, since it is far more likely a config mistake than a wish never to alert; null means the workbook rules.

## Errors

- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X` - `minEvents must be null or a whole number of at least 0, received X` - `windowSeconds must be at least 1, received X` - `duplicate counts for a 3600-second window` - the count errors of `monitor.burn-rate` (`badEvents must not exceed totalEvents: B > T`, ...) - `rules must not be empty; pass null for the SRE workbook rules` - `severity must not be empty` - `shortWindowSeconds must be at least 1, received X` - `shortWindowSeconds must not exceed longWindowSeconds: S > L` - `burnRateMilli must be at least 1, received X` - `no counts for a 3600-second window`

## Sources

Google SRE Workbook, chapter 5 "Alerting on SLOs", the multiwindow, multi-burn-rate alerts section and Table 5-8 (https://sre.google/workbook/alerting-on-slos/): page 1h / 5m / 14.4 / 2%; page 6h / 30m / 6 / 5%; ticket 3d / 6h / 1 / 10%; "make the short window 1/12 the duration of the long window"; and the "Low-Traffic Services and Error Budget Alerting" section of the same chapter for the small-sample problem `minEvents` guards against.

Files

PathBytes
README.md5,115
data/sre-workbook.json1,025
impl/python.py4,963
impl/rust.rs7,231
impl/typescript.ts4,577
vectors.json18,947