Field notes
Error budgets that integration teams will actually follow
Availability targets only help if product, platform, and partner owners share the same burn-rate language.
Many teams inherit a 99.9% target with no shared idea of what burns the remaining 0.1%. Integration work makes this worse: a partner outage, a malformed payload, and a deploy race can all look like “your” downtime depending on who is asked.
An error budget workshop starts with classification. Separate self-inflicted errors from partner-attributable ones, then decide which count against the budget used for release decisions. Clarity here prevents endless post-incident blame.
Burn rate is the operational signal. A slow trickle of 500s over a month may be tolerable; the same volume compressed into a morning is not. Define fast-burn and slow-burn cues so on-call and product leads know when to freeze non-critical releases.
Write the policy on one page. Long SRE handbooks gather dust. A short table — target, budget window, burn cues, freeze rule, who decides — travels better across departments.
In Korea-based delivery teams we often see reliability ownership split between a central platform group and business units that own partner contracts. The workshop format exists precisely to put those groups in the same room with the same numbers.