Functional Weave
Code in Rust

monitor.burn-rate@1.0.0

README.md

1,926 bytes · view raw

# monitor.burn-rate

How fast a service is spending its error budget: the error rate in a window
divided by the error rate its SLO allows, as thousandths (`1000` = 1x).

    burn = (bad / total) / ((10000 - target) / 10000)

At 1x the budget lasts exactly the SLO period (30 days in the workbook); at
14.4x a 30-day budget is gone in 50 hours, and one hour at that rate spends 2%
of it (14.4 x 1h / 720h). Budget spent in a window = burn x window / period.

Burn rate is what SLO alerts compare against a threshold: see
`monitor.burn-rate-alert` for the workbook's multiwindow, multi-burn-rate
rules. On its own it is the number for a dashboard ("burning at 3.2x").

## Decisions

- **Thousandths, half up.** 14.4x is `14400`. Integers keep the three
  languages identical. An alert that must not fire early on rounding should
  compare exactly, as `monitor.burn-rate-alert` does.
- **An empty window is 0**, not an error: no traffic spends no budget, and a
  quiet night should not break the alerting.
- **Exact arithmetic.** bad x 10000 x 1000 passes 2^53 below a billion bad
  events, so the division is exact (BigInt in TypeScript, i128 in Rust,
  Python's integers). The design pinned `math.round-div`; it is not used,
  because its integers stop at 2^53.
- **It works on seconds too**: total = window seconds, bad = down seconds.

## Errors

- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X`
- `totalEvents must be a whole number of at least 0, received X`
- `badEvents must be a whole number of at least 0, received X`
- `badEvents must not exceed totalEvents: B > T`

## Sources

Google SRE Workbook, chapter 5 "Alerting on SLOs", section "Alert on Burn
Rate" (https://sre.google/workbook/alerting-on-slos/): burn rate is how fast,
relative to the SLO, the service consumes the error budget; 1 consumes all of
it in the 30-day window; Table 5-8 uses 14.4, 6 and 1.