# monitor.burn-rate-alert
SLO alerting the way the Google SRE workbook recommends: multiwindow,
multi-burn-rate. Pass the SLO target, the event counts you have for each
window length, and the rules (or null for the workbook's), and it says which
rules fire and at what severity. An optional minimum number of events keeps
a rule quiet on a sample too small to mean anything.
A rule fires when **both** its long window and its short window burn at or
above its threshold (see `monitor.burn-rate`). The long window proves enough
budget has gone to matter; the short window proves it is still going, so the
alert stops soon after the problem does instead of paging for an hour after a
five-minute outage.
## The default rules (rules = null)
Shipped as `data/sre-workbook.json`, from Table 5-8 of the workbook's
"Alerting on SLOs" chapter, for a 30-day SLO period:
| Severity | Long window | Short window | Burn rate | Budget consumed |
|----------|-------------|--------------|-----------|-----------------|
| page | 1 hour | 5 minutes | 14.4 | 2% |
| page | 6 hours | 30 minutes | 6 | 5% |
| ticket | 3 days | 6 hours | 1 | 10% |
Budget consumed = burn rate x long window / 30 days; the short window is 1/12
of the long one, as the workbook suggests. So the windows to count are 300,
1800, 3600, 21600 and 259200 seconds. The table has no `effective` dates: it
is a published recommendation, not a rule that changes in force.
## Small samples: minEvents
Burn rates are ratios, and a ratio of small counts is noise. A new target
checked every 30 s has 11 checks after five minutes; if 2 of them failed, its
error rate is 18%, which burns a 99.5% budget at 36x and pages under every
rule in the table, though nothing much has happened. The workbook names the
problem in its section on low-traffic services ("Low-Traffic Services and
Error Budget Alerting"): one failure among very few requests can look like a
huge burn.
`minEvents` is the usual guard: a rule fires only when its **long** window
saw at least that many events. Its `enoughEvents` says whether it did. The
short window is not guarded, because it is only there to prove the burn is
still happening, and a 5-minute window of a target checked every 30 s never
holds more than 10 checks. Since each rule has its own long window, the guard
is per rule: with 90 events in the last hour and 500 in the last six, a
minimum of 100 silences the 1-hour rule and leaves the 6-hour one free to
page.
`null` or `0` means no minimum, which is what 1.0.0 did: every 1.0.0 vector
gives the same answer here with `minEvents` null (and `enoughEvents` true).
## Changes from 1.0.0
- A fourth parameter, `minEvents: int?`.
- `BurnRuleResult.enoughEvents`.
Both break a caller written for 1.0.0 (one more argument to pass, one more
field in the result), hence 2.0.0 rather than 1.1.0: a project on `^1` keeps
1.0.0 until it moves to `^2`. With `minEvents` null the answers are 1.0.0's.
## Decisions
- **Firing is exact.** It compares bad x 10,000,000 against threshold x total x
(10000 - target), not the rounded `longBurnMilli`: a burn of 14.3995x shows
as `14400` but does not fire a 14.4x rule.
- **Severity is the first firing rule in rule order.** List pages before
tickets (the workbook table already does), and a fast burn that also trips
the ticket rule still pages.
- **An empty window (total 0) never fires.**
- **Every window is checked**, including ones no rule uses: bad counts in any
of them are a bug upstream.
- Each rule needs counts for exactly its two window lengths, so the caller
counts once per distinct length; a rule with short = long is allowed (a
single-window alert).
- `minEvents` guards the long window only, and counts every event in it,
good or bad.
- `rules` = `[]` is an error, since it is far more likely a config mistake
than a wish never to alert; null means the workbook rules.
## Errors
- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X`
- `minEvents must be null or a whole number of at least 0, received X`
- `windowSeconds must be at least 1, received X`
- `duplicate counts for a 3600-second window`
- the count errors of `monitor.burn-rate` (`badEvents must not exceed totalEvents: B > T`, ...)
- `rules must not be empty; pass null for the SRE workbook rules`
- `severity must not be empty`
- `shortWindowSeconds must be at least 1, received X`
- `shortWindowSeconds must not exceed longWindowSeconds: S > L`
- `burnRateMilli must be at least 1, received X`
- `no counts for a 3600-second window`
## Sources
Google SRE Workbook, chapter 5 "Alerting on SLOs", the multiwindow,
multi-burn-rate alerts section and Table 5-8
(https://sre.google/workbook/alerting-on-slos/):
page 1h / 5m / 14.4 / 2%; page 6h / 30m / 6 / 5%; ticket 3d / 6h / 1 / 10%;
"make the short window 1/12 the duration of the long window"; and the
"Low-Traffic Services and Error Budget Alerting" section of the same chapter
for the small-sample problem `minEvents` guards against.