Functional Weave
Code in TypeScript

monitor.burn-rate-alert@1.0.0

README.md

3,210 bytes · view raw

# monitor.burn-rate-alert

SLO alerting the way the Google SRE workbook recommends: multiwindow,
multi-burn-rate. Pass the SLO target, the event counts you have for each
window length, and the rules (or null for the workbook's), and it says which
rules fire and at what severity.

A rule fires when **both** its long window and its short window burn at or
above its threshold (see `monitor.burn-rate`). The long window proves enough
budget has gone to matter; the short window proves it is still going, so the
alert stops soon after the problem does instead of paging for an hour after a
five-minute outage.

## The default rules (rules = null)

Shipped as `data/sre-workbook.json`, from Table 5-8 of the workbook's
"Alerting on SLOs" chapter, for a 30-day SLO period:

| Severity | Long window | Short window | Burn rate | Budget consumed |
|----------|-------------|--------------|-----------|-----------------|
| page     | 1 hour      | 5 minutes    | 14.4      | 2%              |
| page     | 6 hours     | 30 minutes   | 6         | 5%              |
| ticket   | 3 days      | 6 hours      | 1         | 10%             |

Budget consumed = burn rate x long window / 30 days; the short window is 1/12
of the long one, as the workbook suggests. So the windows to count are 300,
1800, 3600, 21600 and 259200 seconds. The table has no `effective` dates: it
is a published recommendation, not a rule that changes in force.

## Decisions

- **Firing is exact.** It compares bad x 10,000,000 against threshold x total x
  (10000 - target), not the rounded `longBurnMilli`: a burn of 14.3995x shows
  as `14400` but does not fire a 14.4x rule.
- **Severity is the first firing rule in rule order.** List pages before
  tickets (the workbook table already does), and a fast burn that also trips
  the ticket rule still pages.
- **An empty window (total 0) never fires.**
- **Every window is checked**, including ones no rule uses: bad counts in any
  of them are a bug upstream.
- Each rule needs counts for exactly its two window lengths, so the caller
  counts once per distinct length; a rule with short = long is allowed (a
  single-window alert).
- `rules` = `[]` is an error, since it is far more likely a config mistake
  than a wish never to alert; null means the workbook rules.

## Errors

- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X`
- `windowSeconds must be at least 1, received X`
- `duplicate counts for a 3600-second window`
- the count errors of `monitor.burn-rate` (`badEvents must not exceed totalEvents: B > T`, ...)
- `rules must not be empty; pass null for the SRE workbook rules`
- `severity must not be empty`
- `shortWindowSeconds must be at least 1, received X`
- `shortWindowSeconds must not exceed longWindowSeconds: S > L`
- `burnRateMilli must be at least 1, received X`
- `no counts for a 3600-second window`

## Sources

Google SRE Workbook, chapter 5 "Alerting on SLOs", the multiwindow,
multi-burn-rate alerts section and Table 5-8
(https://sre.google/workbook/alerting-on-slos/):
page 1h / 5m / 14.4 / 2%; page 6h / 30m / 6 / 5%; ticket 3d / 6h / 1 / 10%;
"make the short window 1/12 the duration of the long window".