Principle
Let reliability data slow feature velocity when the budget is spent
Reliability work needs permission to outrank new features before the next outage makes the choice for you.
When it fits
- A team keeps shipping changes while user-facing reliability is already below the agreed target.
When to avoid it
- An arbitrary SLO creates arbitrary behavior; error budgets only help when the objective reflects real user value and risk.
Why it matters
Define a user-facing service objective and an allowed error budget for a period. When reliability is healthy, change can proceed normally. When the budget is exhausted or SLO misses persist, shift capacity from new changes toward reliability until the service returns to the agreed range. Treat the policy as prioritization, not punishment.
An example
Pause nonessential agent-feature rollouts when repeated tool failures consume the monthly success-rate budget and focus on reliability fixes.
Check your result
Reliability deterioration changes work priority through an agreed rule rather than ad hoc escalation.
Keep this limit in mind
- An arbitrary SLO creates arbitrary behavior; error budgets only help when the objective reflects real user value and risk.
Evidence and sources
Google's example error-budget policy uses SLO performance to decide when teams should continue releases versus focus effort on reliability.
Error-budget policy is an organizational control model and requires an SLO that reflects user value.
Example Error Budget Policy · Goals and SLO miss policy