# monitor.error-budget
The error budget of a service level objective (SLO) over one window: how many
bad events the SLO allows, how many are left, and what share of the budget is
spent. A 99.9% SLO (`9990`) over a million requests allows 1000 failures; 250
failures have spent 25% (`2500` basis points) and leave 750.
## Events or seconds
It works for request-based SLOs (total requests, failed requests) and for
time-based ones (seconds in the window, seconds down). 99.9% over 30 days is
2,592,000 seconds with a budget of **2592 seconds, 43m 12s**. The `checks`,
`downChecks` and `downSeconds` from `monitor.uptime` are the counts to pass.
## The numbers
- `allowedBad` = floor(total x (10000 - target) / 10000): whole events, so it
floors. A budget of 1.5 events allows 1.
- `consumedBasisPoints` = bad / exact budget, in basis points, rounded half up.
It is measured against the *exact* budget, not the floored one, so 1 bad
event out of a 1.5-event budget is 6667 (66.67%), not 10000. It exceeds
10000 once the budget is overspent; every request failing at 99.9% is
10,000,000 (1000x).
- `remainingBad` = allowedBad - bad and `remainingBasisPoints` = 10000 -
consumed. Both go negative when over budget, which is how far over.
- `exhausted` is exact: bad x 10000 > total x (10000 - target). Exactly at
budget is fully spent (`consumed` 10000, `remaining` 0) but not exhausted;
one more bad event is: overspent, the point where an error budget policy
usually freezes releases.
- No events yet (`total` 0) spends nothing: consumed 0, not exhausted.
The design pinned `math.round-div` for the rounding; it is not used, because
bad x 10000 x 10000 passes 2^53 (the limit of `round-div`'s integers) at only
90 million bad events. The division is done exactly instead, with BigInt in
TypeScript, i128 in Rust and Python's own integers.
## Errors
- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X`
- `totalEvents must be a whole number of at least 0, received X`
- `badEvents must be a whole number of at least 0, received X`
- `badEvents must not exceed totalEvents: B > T`
## Sources
- Google SRE Workbook, "Implementing SLOs" and "Alerting on SLOs"
(https://sre.google/workbook/alerting-on-slos/): error budget = 1 - SLO,
measured over a 30-day window.
- Google SRE Book, "Embracing Risk", error budgets
(https://sre.google/sre-book/embracing-risk/).