# monitor.burn-rate-alert SLO alerting the way the Google SRE workbook recommends: multiwindow, multi-burn-rate. Pass the SLO target, the event counts you have for each window length, and the rules (or null for the workbook's), and it says which rules fire and at what severity. An optional minimum number of events keeps a rule quiet on a sample too small to mean anything. A rule fires when **both** its long window and its short window burn at or above its threshold (see `monitor.burn-rate`). The long window proves enough budget has gone to matter; the short window proves it is still going, so the alert stops soon after the problem does instead of paging for an hour after a five-minute outage. ## The default rules (rules = null) Shipped as `data/sre-workbook.json`, from Table 5-8 of the workbook's "Alerting on SLOs" chapter, for a 30-day SLO period: | Severity | Long window | Short window | Burn rate | Budget consumed | |----------|-------------|--------------|-----------|-----------------| | page | 1 hour | 5 minutes | 14.4 | 2% | | page | 6 hours | 30 minutes | 6 | 5% | | ticket | 3 days | 6 hours | 1 | 10% | Budget consumed = burn rate x long window / 30 days; the short window is 1/12 of the long one, as the workbook suggests. So the windows to count are 300, 1800, 3600, 21600 and 259200 seconds. The table has no `effective` dates: it is a published recommendation, not a rule that changes in force. ## Small samples: minEvents Burn rates are ratios, and a ratio of small counts is noise. A new target checked every 30 s has 11 checks after five minutes; if 2 of them failed, its error rate is 18%, which burns a 99.5% budget at 36x and pages under every rule in the table, though nothing much has happened. The workbook names the problem in its section on low-traffic services ("Low-Traffic Services and Error Budget Alerting"): one failure among very few requests can look like a huge burn. `minEvents` is the usual guard: a rule fires only when its **long** window saw at least that many events. Its `enoughEvents` says whether it did. The short window is not guarded, because it is only there to prove the burn is still happening, and a 5-minute window of a target checked every 30 s never holds more than 10 checks. Since each rule has its own long window, the guard is per rule: with 90 events in the last hour and 500 in the last six, a minimum of 100 silences the 1-hour rule and leaves the 6-hour one free to page. `null` or `0` means no minimum, which is what 1.0.0 did: every 1.0.0 vector gives the same answer here with `minEvents` null (and `enoughEvents` true). ## Changes from 1.0.0 - A fourth parameter, `minEvents: int? = null`, which a caller may leave out. - `BurnRuleResult.enoughEvents: bool = true`, the last field. Both are additions with a default, so this is a minor version: a call written for 1.0.0 still compiles and gives 1.0.0's answer, and every 1.0.0 vector is carried over unedited (a result it describes without `enoughEvents` means `true`, the default). In each language: - TypeScript: `burnRateAlert(target, windows, rules)` as before, or `burnRateAlert(target, windows, rules, 20)`. - Python: `burn_rate_alert(target, windows, rules)`, or `min_events=20`. - Rust, which has no default arguments: `burn_rate_alert(target, &windows, rules)` as before, or `burn_rate_alert_with(target, &windows, rules, BurnRateAlertOptions { min_events: Some(20), ..Default::default() })`. Write the `..Default::default()`: a later minor version may add an option. 2.0.0 is the same function with `minEvents` required (and `enoughEvents` before `firing`), published before fune had defaults. It stays published, and correct, for the projects already on `^2`; new work happens on this line, so a new project should use `^1.1`. ## Decisions - **Firing is exact.** It compares bad x 10,000,000 against threshold x total x (10000 - target), not the rounded `longBurnMilli`: a burn of 14.3995x shows as `14400` but does not fire a 14.4x rule. - **Severity is the first firing rule in rule order.** List pages before tickets (the workbook table already does), and a fast burn that also trips the ticket rule still pages. - **An empty window (total 0) never fires.** - **Every window is checked**, including ones no rule uses: bad counts in any of them are a bug upstream. - Each rule needs counts for exactly its two window lengths, so the caller counts once per distinct length; a rule with short = long is allowed (a single-window alert). - `minEvents` guards the long window only, and counts every event in it, good or bad. - `rules` = `[]` is an error, since it is far more likely a config mistake than a wish never to alert; null means the workbook rules. ## Errors - `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X` - `minEvents must be null or a whole number of at least 0, received X` - `windowSeconds must be at least 1, received X` - `duplicate counts for a 3600-second window` - the count errors of `monitor.burn-rate` (`badEvents must not exceed totalEvents: B > T`, ...) - `rules must not be empty; pass null for the SRE workbook rules` - `severity must not be empty` - `shortWindowSeconds must be at least 1, received X` - `shortWindowSeconds must not exceed longWindowSeconds: S > L` - `burnRateMilli must be at least 1, received X` - `no counts for a 3600-second window` ## Sources Google SRE Workbook, chapter 5 "Alerting on SLOs", the multiwindow, multi-burn-rate alerts section and Table 5-8 (https://sre.google/workbook/alerting-on-slos/): page 1h / 5m / 14.4 / 2%; page 6h / 30m / 6 / 5%; ticket 3d / 6h / 1 / 10%; "make the short window 1/12 the duration of the long window"; and the "Low-Traffic Services and Error Budget Alerting" section of the same chapter for the small-sample problem `minEvents` guards against.