monitor.error-budget
How much of an SLO's error budget is allowed, used and left, from good and bad event counts or seconds.
1.0.0 · published 2026-10-03 by charlie · Anterra
Pinned by 19 tests, run in TypeScript, Python and Rust.
What it does
The error budget of a service level objective (SLO) over one window: how many bad events the SLO allows, how many are left, and what share of the budget is spent. A 99.9% SLO (`9990`) over a million requests allows 1000 failures; 250 failures have spent 25% (`2500` basis points) and leave 750.
## Events or seconds
For example
error_budget(99.9%, 1,000,000, 250)→ allowed bad 1,000, remaining bad 750, consumed basis points 25%, remaining basis points 75%, exhausted false 99.9% of a million requests allows 1000 bad; 250 spent is a quarter of the budgeterror_budget(99.9%, 2,592,000, 0)→ allowed bad 2,592, remaining bad 2,592, consumed basis points 0%, remaining basis points 100%, exhausted false a time budget: 99.9% over 30 days allows 2592 down seconds (43m 12s)error_budget(99.9%, 2,592,000, 1,296)→ allowed bad 2,592, remaining bad 1,296, consumed basis points 50%, remaining basis points 50%, exhausted false half the 30-day time budget spent after 1296 down seconds
The function
The same function in TypeScript, Python and Rust, pinned by the same tests. Pick your language; the choice follows you around the registry.
def error_budget(target_basis_points: int, total_events: int, bad_events: int) -> ErrorBudget
| target_basis_points | int | the SLO: 9990 = 99.9%; 1 to 9999 |
| total_events | int | requests (or seconds) in the SLO window |
| bad_events | int | the failed ones (or down seconds); at most totalEvents |
| returns | ErrorBudget |
The type it declares, generated into your project
@dataclass(frozen=True)
class ErrorBudget:
"""The budget in events and in basis points of itself."""
#: floor(total x (10000 - target) / 10000)
allowed_bad: int
#: allowedBad - bad; negative once over budget
remaining_bad: int
#: bad as a share of the exact budget, half-up; over 10000 once over budget
consumed_basis_points: int
#: 10000 - consumedBasisPoints; may be negative
remaining_basis_points: int
#: bad is strictly more than the exact budget
exhausted: bool
Your code names it in one line, in the file that uses it
from fune.monitor.error_budget import error_budget # monitor.error-budget@^1
from .monitor_error_budget_types import ErrorBudget
def _whole(value: object) -> bool:
return isinstance(value, int) and not isinstance(value, bool)
def _check_counts(target_basis_points: int, total_events: int, bad_events: int) -> None:
if not _whole(target_basis_points) or target_basis_points < 1 or target_basis_points > 9999:
raise ValueError("targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received %s" % (target_basis_points,))
if not _whole(total_events) or total_events < 0:
raise ValueError("totalEvents must be a whole number of at least 0, received %s" % (total_events,))
if not _whole(bad_events) or bad_events < 0:
raise ValueError("badEvents must be a whole number of at least 0, received %s" % (bad_events,))
if bad_events > total_events:
raise ValueError("badEvents must not exceed totalEvents: %d > %d" % (bad_events, total_events))
def _half_up(n: int, d: int) -> int:
# Non-negative operands only, so half-up needs no sign handling.
return (2 * n + d) // (2 * d)
def error_budget(target_basis_points: int, total_events: int, bad_events: int) -> ErrorBudget:
"""The error budget of an SLO over one window: how many bad events it
allows, how many are left, and how much of it is spent. Exact integers
throughout."""
_check_counts(target_basis_points, total_events, bad_events)
# The exact budget is budget_times_10000 / 10000 events; keep it as a fraction.
budget_times_10000 = total_events * (10000 - target_basis_points)
allowed_bad = budget_times_10000 // 10000
consumed = 0 if total_events == 0 else _half_up(bad_events * 10000 * 10000, budget_times_10000)
return ErrorBudget(
allowed_bad=allowed_bad,
remaining_bad=allowed_bad - bad_events,
consumed_basis_points=consumed,
remaining_basis_points=10000 - consumed,
exhausted=bad_events * 10000 > budget_times_10000,
)Install
fune build
With that line in your source, in a Python project (language python in fune.project), fune build resolves it and nothing else, pins them in fune.lock, downloads only the Python package of each, and builds the code above into your project’s .fune/build, one readable file per capability with a header linking back here. Or pin a range in fune.project and build in one step:
fune add monitor.error-budget
The manifest, vectors and README with only the Python implementation. Install it without the registry with fune add ./monitor.error-budget-1.0.0-python.fune, or fetch it from a terminal with fune pull monitor.error-budget@1.0.0:python.
The whole function, every language, is one file too: monitor.error-budget-1.0.0.fune, 16,523 bytes, sha256 0895adf982e895f32250a3b40eac437e9a220efbad181a65b878c3c16c09fc1b. It installs into a project of any language.
Customise it in your app
The seams this capability offers. Put a marker directly above a function of your own and fune build wires it into the built code; the package on the registry is not changed, the built file’s header lists it under CUSTOMISED, and fune hooks lists every hook in the project. How hooks work.
before — your function gets the arguments and returns them, changed or not, or throws to refuse the call.
# fune: before monitor.error-budget
after — your function gets the result and the arguments, and returns the final result.
# fune: after monitor.error-budget
replace — it requires no other capability, so there is no dependency to replace.
step — your function runs at a numbered point inside the function’s body, receives the in-scope values it names as parameters, and may return replacements. List the points with fune show monitor.error-budget --steps.
# fune: step monitor.error-budget after <n|label>
Tests
A version published now needs at least 8 tests for every function, and one that expects the error for each function that throws; the registry refuses it otherwise. fune verify --all runs each case in TypeScript, Python and Rust, and a project runs them again with fune verify. This page lists the cases; it does not run them. The exact JSON is vectors.json.
| Case | Arguments | Expected | |
|---|---|---|---|
| 99.9% of a million requests allows 1000 bad; 250 spent is a quarter of the budget | 99.9%, 1,000,000, 250 | → | allowed bad 1,000, remaining bad 750, consumed basis points 25%, remaining basis points 75%, exhausted false |
| a time budget: 99.9% over 30 days allows 2592 down seconds (43m 12s) | 99.9%, 2,592,000, 0 | → | allowed bad 2,592, remaining bad 2,592, consumed basis points 0%, remaining basis points 100%, exhausted false |
| half the 30-day time budget spent after 1296 down seconds | 99.9%, 2,592,000, 1,296 | → | allowed bad 2,592, remaining bad 1,296, consumed basis points 50%, remaining basis points 50%, exhausted false |
| exactly at budget is fully spent but not exhausted | 99.9%, 1,000,000, 1,000 | → | allowed bad 1,000, remaining bad 0, consumed basis points 100%, remaining basis points 0%, exhausted false |
| one bad event over budget is exhausted, and remaining goes negative | 99.9%, 1,000,000, 1,001 | → | allowed bad 1,000, remaining bad -1, consumed basis points 100.1%, remaining basis points -0.1%, exhausted true |
| no events yet: nothing allowed, nothing spent | 99.9%, 0, 0 | → | allowed bad 0, remaining bad 0, consumed basis points 0%, remaining basis points 100%, exhausted false |
| a budget of 1.5 events: allowedBad floors to 1, but 1 bad is only two thirds spent and not exhausted | 99.9%, 1,500, 1 | → | allowed bad 1, remaining bad 0, consumed basis points 66.67%, remaining basis points 33.33%, exhausted false |
| a budget of 1.5 events with 2 bad is exhausted | 99.9%, 1,500, 2 | → | allowed bad 1, remaining bad -1, consumed basis points 133.33%, remaining basis points -33.33%, exhausted true |
| consumed rounds half up: 2.5 basis points becomes 3 | 90%, 40,000, 1 | → | allowed bad 4,000, remaining bad 3,999, consumed basis points 0.03%, remaining basis points 99.97%, exhausted false |
| consumed rounds half up: 62.5 basis points becomes 63 | 99.9%, 160,000, 1 | → | allowed bad 160, remaining bad 159, consumed basis points 0.63%, remaining basis points 99.37%, exhausted false |
Show the other 9 tests
| Case | Arguments | Expected | |
|---|---|---|---|
| every request failing at 99.9% burns the budget a thousand times over | 99.9%, 1,000, 1,000 | → | allowed bad 1, remaining bad -999, consumed basis points 100000%, remaining basis points -99900%, exhausted true |
| the loosest target, 0.01%, allows 9999 of 10000 bad | 0.01%, 10,000, 10,000 | → | allowed bad 9,999, remaining bad -1, consumed basis points 100.01%, remaining basis points -0.01%, exhausted true |
| a trillion requests: products pass 2^53 and stay exact | 99.9%, 1,000,000,000,000, 400,049,999 | → | allowed bad 1,000,000,000, remaining bad 599,950,001, consumed basis points 40%, remaining basis points 60%, exhausted false |
| a 100% target has no budget and is an error | 100%, 100, 0 | → | error: targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received 10000 |
| a 0% target is an error | 0%, 100, 0 | → | error: targetBasisPoints must be a whole number from 1 to 9999 |
| negative total is an error | 99.9%, -1, 0 | → | error: totalEvents must be a whole number of at least 0, received -1 |
| negative bad is an error | 99.9%, 10, -1 | → | error: badEvents must be a whole number of at least 0, received -1 |
| more bad events than events is an error | 99.9%, 4, 5 | → | error: badEvents must not exceed totalEvents: 5 > 4 |
| a fractional count is an error | 99.9%, 2.5, 0 | → | error: totalEvents must be a whole number of at least 0, received 2.5 |
More from the author
It works for request-based SLOs (total requests, failed requests) and for time-based ones (seconds in the window, seconds down). 99.9% over 30 days is 2,592,000 seconds with a budget of **2592 seconds, 43m 12s**. The `checks`, `downChecks` and `downSeconds` from `monitor.uptime` are the counts to pass.
## The numbers
- `allowedBad` = floor(total x (10000 - target) / 10000): whole events, so it floors. A budget of 1.5 events allows 1. - `consumedBasisPoints` = bad / exact budget, in basis points, rounded half up. It is measured against the *exact* budget, not the floored one, so 1 bad event out of a 1.5-event budget is 6667 (66.67%), not 10000. It exceeds 10000 once the budget is overspent; every request failing at 99.9% is 10,000,000 (1000x). - `remainingBad` = allowedBad - bad and `remainingBasisPoints` = 10000 - consumed. Both go negative when over budget, which is how far over. - `exhausted` is exact: bad x 10000 > total x (10000 - target). Exactly at budget is fully spent (`consumed` 10000, `remaining` 0) but not exhausted; one more bad event is: overspent, the point where an error budget policy usually freezes releases. - No events yet (`total` 0) spends nothing: consumed 0, not exhausted.
The design pinned `math.round-div` for the rounding; it is not used, because bad x 10000 x 10000 passes 2^53 (the limit of `round-div`'s integers) at only 90 million bad events. The division is done exactly instead, with BigInt in TypeScript, i128 in Rust and Python's own integers.
## Errors
- `targetBasisPoints must be a whole number from 1 to 9999 (10000 leaves no error budget), received X` - `totalEvents must be a whole number of at least 0, received X` - `badEvents must be a whole number of at least 0, received X` - `badEvents must not exceed totalEvents: B > T`
## Sources
- Google SRE Workbook, "Implementing SLOs" and "Alerting on SLOs" (https://sre.google/workbook/alerting-on-slos/): error budget = 1 - SLO, measured over a 30-day window. - Google SRE Book, "Embracing Risk", error budgets (https://sre.google/sre-book/embracing-risk/).
Files
| Path | Bytes |
|---|---|
| README.md | 2,439 |
| impl/python.py | 1,979 |
| impl/rust.rs | 2,988 |
| impl/typescript.ts | 2,040 |
| vectors.json | 3,997 |