Skip to lesson
Exit
IaC, Cost Engineering & Platform Unification1 / 2

1 min lesson

Error budgets are the dial

Rebuild the parts of "Error budgets are the dial", then say why each one matters.

Step 1 of 2

Error budgets are the dialspending reliability you've already promised

An error budget turns reliability from a vibe into a number you can spend. If the SLO is 99.9% and you're consistently hitting 99.99%, you are over-delivering - and over-delivery often means over-paying for redundancy nobody asked for. That gap is a budget you can spend on cost cuts or velocity.

Reading the error budget
Burning budget fast
Buy reliability: add redundancy, slow risky changes, invest in resilience
Hitting SLO with margin
Hold - you're priced about right; don't gold-plate further
Consistently far above SLO
Spend the surplus: remove redundancy that isn't paying for itself, take more velocity
No SLO at all
You can't reason about the tradeoff yet; define the SLI/SLO first

The error budget tells you which direction to move on the cost/reliability dial.

Removing redundancy feels dangerous, so make it evidence-led. If a service has held three nines for a year against a 99.9% SLO, dropping from a hot standby in a third region to a warm one is spending surplus budget, not gambling. State it that way and the move stops sounding reckless.