An error budget turns 'is it reliable enough' into a number
99.9 percent over 30 days is 43 minutes of failure. Spend it on incidents and there is none left for risky deploys. Spend none and you are being more careful than the users asked for.
Track the budget burn on a graph. When it is running out, the team ships reliability work instead of features, by agreement made in advance rather than in an argument during the incident.
observabilitysre