# monitor.error-budget The error budget of a service level objective (SLO) over one window: how many bad events the SLO allows, how many are left, and what share of the budget is spent. A 99.9% SLO (`9990`) over a million requests allows 1000 failures; 250 failures have spent 25% (`2500` basis points) and leave 750. ## Events or seconds It works for request-based SLOs (total requests, failed requests) and for time-based ones (seconds in the window, seconds down). 99.9% over 30 days is 2,592,000 seconds with a budget of **2592 seconds, 43m 12s**. The `checks`, `downChecks` and `downSeconds` from `monitor.uptime` are the counts to pass. ## The numbers - `allowedBad` = floor(total x (10000 - target) / 10000): whole events, so it floors. A budget of 1.5 events allows 1. - `consumedBasisPoints` = bad / exact budget, in basis points, rounded half up. It is measured against the *exact* budget, not the floored one, so 1 bad event out of a 1.5-event budget is 6667 (66.67%), not 10000. It exceeds 10000 once the budget is overspent; every request failing at 99.9% is 10,000,000 (1000x). - `remainingBad` = allowedBad - bad and `remainingBasisPoints` = 10000 - consumed. Both go negative when over budget, which is how far over. - `exhausted` is exact: bad x 10000 > total x (10000 - target). Exactly at budget is fully spent (`consumed` 10000, `remaining` 0) but not exhausted; one more bad event is: overspent, the point where an error budget policy usually freezes releases. - No events yet (`total` 0) spends nothing: consumed 0, not exhausted. The design pinned `math.round-div` for the rounding; it is not used, because bad x 10000 x 10000 passes 2^53 (the limit of `round-div`'s integers) at only 90 million bad events. The division is done exactly instead, with BigInt in TypeScript, i128 in Rust and Python's own integers. ## Errors - `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X` - `totalEvents must be a whole number of at least 0, received X` - `badEvents must be a whole number of at least 0, received X` - `badEvents must not exceed totalEvents: B > T` ## Sources - Google SRE Workbook, "Implementing SLOs" and "Alerting on SLOs" (https://sre.google/workbook/alerting-on-slos/): error budget = 1 - SLO, measured over a 30-day window. - Google SRE Book, "Embracing Risk", error budgets (https://sre.google/sre-book/embracing-risk/).