Skip to main content
DevOps & Cloud

What Two Years of On-Call Taught Me About SLOs

G
Grace Wanjiru

Aug 8, 2026 · 1 min read · 3 comments

Our pager used to go off for CPU spikes nobody cared about. Customers noticed outages before we did. SLOs fixed both.

Start from the user

Pick one journey ("a customer can complete checkout") and measure its success rate and latency at the edge, not on one pod.

Alert on burn rate

An error budget turns "is this bad?" into arithmetic. We page on fast burn and open a ticket on slow burn.

Review monthly

If you never spend your error budget, your target is too loose. If you always overspend it, stop shipping features and fix reliability.

The best part: on-call is quiet now, and when the pager rings it matters.

Share WhatsApp Post LinkedIn
4 Like 0 Insightful 1 Fire
Sign in to react and save

Discussion (3)

Sign in to join the discussion.

Ibrahim Musa ·

How did you pick the first journey to measure? We're stuck debating it.

Grace Wanjiru Author ·

@ibrahim-musa we picked the journey that generated the most support tickets. Easy win.

Kofi Asante ·

Burn-rate alerts changed our on-call completely. Great write-up.