Building an On-Call Rotation That Does Not Burn People Out

The pager went off at 3:42 on a Tuesday morning. A payment reconciliation job had stalled, and the alert routed to an engineer who had been on call for eleven...

Originally published onanselmfowel.com

The pager went off at 3:42 on a Tuesday morning. A payment reconciliation job had stalled, and the alert routed to an engineer who had been on call for eleven straight days because the person after him had quit and we hadn't backfilled the rotation. He fixed it in twenty minutes. Then he sent me a resignation email at 9 a.m. I don't blame him. I blame the rotation we built, which was less a schedule and more a slow-motion attrition machine.

Building an On-Call Rotation That Does Not Burn People Out
Building an On-Call Rotation That Does Not Burn People Out

On-call is one of those things every engineering org claims to care about and very few actually design. Most rotations are inherited, not chosen. Someone set them up three years ago, the team doubled, the alert volume tripled, and nobody went back to ask whether the thing still made sense. I want to talk about what actually keeps people from quitting, because I have gotten this badly wrong and slowly less wrong over about a decade of running teams that keep money moving.

The real cost is not the night, it is the anticipation

People assume the damage from on-call comes from getting woken up. It doesn't, mostly. In a healthy rotation you get paged maybe once or twice a week, and a lot of that is during the day. The corrosive part is the low-grade dread that colors the entire week you're holding the pager. You don't drink the second glass of wine. You don't go to the movie where you'd lose signal. You keep the laptop in the bag at the kid's birthday. The night you actually get paged is almost a relief compared to the six evenings you spent bracing for it.

Once I understood that, my whole approach changed. The goal isn't to make the 3 a.m. page pleasant, because it never will be. The goal is to shrink the anxiety footprint of the week. That means predictability, a hard cap on how much can go wrong, and the confidence that when something does break, the runbook and the tooling will carry most of the weight. An engineer who trusts the system relaxes. An engineer who has been burned by a vague, undocumented page at 4 a.m. never fully relaxes again.

Fix your alerts before you touch the schedule

No rotation design survives contact with a noisy alerting setup. If your on-call person gets fifteen pages a night and thirteen of them are non-actionable, the schedule is irrelevant. You are burning people out with garbage, and rearranging who receives the garbage doesn't help.

The first team I truly fixed had an alert-to-action ratio that was frankly embarrassing. We measured it for two weeks: 214 pages, of which 31 required any human action at all. The rest were flapping thresholds, a disk-space warning that self-resolved every night when log rotation ran, and one legendary alert that fired whenever a batch job took longer than usual, which was every single time the batch was large. We spent a full sprint doing nothing but deleting and tuning alerts. Page volume dropped by roughly 80 percent. That sprint did more for morale than any comp adjustment I've ever made.

My rule now is blunt: if an alert pages a human and the correct response is to acknowledge it and go back to sleep, it is not an alert. It is a log line wearing a costume. Demote it. Every page that reaches a person should be actionable, urgent, and real. If it is not all three, it does not get to interrupt someone's evening.

Severity levels that actually mean something

Most severity systems are theater. Everything is a P2 because P1 feels scary and P3 feels like admitting it doesn't matter. The result is that nothing gets the urgency it deserves and everything gets a little of the panic it doesn't.

I insist on severity definitions tied to concrete, unambiguous outcomes, not vibes. For us, in payments, it looks roughly like this:

  • Sev1: money is moving incorrectly, or not moving at all, for real customers right now. Wake anyone. This is the only tier that pages at 3 a.m. without hesitation.
  • Sev2: a customer-facing degradation with a workaround, or a risk that will become Sev1 within hours if ignored. Page during waking hours, escalate if it worsens.
  • Sev3: internal or non-urgent. Goes to a queue, handled next business day. Never pages anyone.

The discipline is in refusing to let things drift upward. A Sev3 does not get to become a Sev2 because someone is anxious about it. And a genuine Sev1 must never, ever sit in a queue because the on-call person wasn't sure it qualified. When people trust that the severity of the page matches the severity of the reality, they stop treating every buzz as a potential catastrophe. That trust is the whole game.

Follow the sun if you possibly can

The single biggest structural improvement I ever made to an on-call rotation was geographic. When we opened a small engineering pod in a timezone eight hours off from our headquarters, we split the rotation so that each region covered its own daylight hours. Nobody got paged at night unless it was a genuine Sev1 that the awake region couldn't handle alone.

Overnight pages dropped to a trickle. The engineers who used to dread their on-call week started treating it as a fairly ordinary week with a bit more focus on operational work. I understand not every company can hire across timezones, and I'm not going to pretend a five-person startup can run follow-the-sun. But if you have the headcount and you're keeping people up at night out of pure inertia, you are choosing that pain. It is a solvable problem for a lot more teams than admit it.

If you genuinely cannot spread across timezones, then be honest about the cost and pay for it, both in money and in time back. Which brings me to the part everyone skips.

Compensate it, and give the time back

On-call is labor. It is a constraint on someone's life during hours they would otherwise own completely. Pretending it is just part of the job, folded invisibly into salary, is how you build quiet resentment that surfaces as attrition eighteen months later.

If you cannot afford to compensate on-call, you cannot afford to run on-call. You are simply financing it with your engineers' goodwill, and that account runs dry faster than you think.

We pay a flat stipend per on-call week, more for the overnight-heavy rotations, and it is not trivial money. Just as important, we have a hard rule that if you get paged after midnight, you do not start before noon the next day, no questions, no Slack passive-aggression about it. If a bad night turns into a bad incident that eats your morning, that time comes back to you. Sleep is not a resource we get to borrow against for free. The stipend costs us real budget, and I have defended it in finance reviews more than once. It is cheaper than replacing a senior engineer, which easily runs into six figures once you count the recruiter, the months of ramp, and the institutional knowledge that walks out the door.

Runbooks, and shrinking the blast radius

The difference between a fifteen-minute page and a three-hour page is almost never the engineer's raw talent. It is whether there is a runbook that tells them what this alert means, what usually causes it, and the first three things to check. A senior engineer with no context is slower than a mid-level engineer with a good runbook, every time.

I hold a hard line: an alert that pages people is not allowed to exist without a linked runbook. If you create the alert, you write the runbook. If the runbook is stale, fixing it is part of resolving the incident. This sounds bureaucratic and I promise you it is the opposite. It is the thing that lets a person handle a page for a system they don't own without spiraling into panic at 4 a.m. It is also how knowledge stops living exclusively in the heads of the two people who wrote the service and are, coincidentally, the two people who are most burned out.

The other half is architectural. Feature flags, circuit breakers, and the ability to degrade gracefully mean that a lot of would-be Sev1s become "flip the flag, go back to bed, fix it properly tomorrow." Every kill switch you build is a night of sleep you are banking for a future colleague. I would rather ship a feature a week late with a clean off switch than ship it on time and hand someone a 3 a.m. problem with no safe way to stop the bleeding.

The handoff is a real meeting, not a Slack message

Rotations fail at the seams. The person coming off call knows things the person coming on does not: the flaky job that's been threatening to fall over, the deploy that went out Friday, the customer escalation that might turn into a page. If that context evaporates because the handoff was a one-line message in a channel, the incoming engineer inherits landmines they can't see.

We do a fifteen-minute live handoff, every rotation, no exceptions. What's currently on fire, what's smoldering, what changed this week, what to watch. It is short and it is boring and it has prevented more repeat incidents than any dashboard I've ever built. Boring handoffs are a sign of health, not a waste of time.

Blameless, or the whole thing rots

Here is where I get opinionated. If your postmortem culture punishes the person who was holding the pager when something broke, you have already lost. People will start under-reporting incidents, resolving things quietly, and avoiding on-call entirely. The rotation becomes a hot potato nobody wants to hold, and your best engineers negotiate their way out of it first.

The engineer on call did not cause the outage. The system that allowed a single bad deploy to take down payments caused the outage. The person was just standing there when the accumulated debt came due. I have sat in postmortems where a director wanted a name and a reprimand, and I have said, flatly, that we will be fixing the process and not the person, and if that is unacceptable we can discuss it after the incident is closed. Blameless is not softness. It is the only way you get honest data about how your systems actually fail.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

The rotation is a product too

Treat your on-call rotation as a product, with your own engineers as the users, and hold it to the same standard as anything a customer touches. That single framing does most of the work, because you would never ship a customer product, watch its error rate climb steadily for three years, and then shrug off the churn as an unavoidable cost of doing business. You'd open a ticket.

The teams that burn out are not unlucky, and they are not weak. They are running a rotation nobody has looked at since it was set up, and they mistake the resulting attrition for the natural price of serious work. It isn't. It is a design failure, and design failures get fixed the moment one person decides to actually look at the thing instead of stepping around it for another quarter.

Chat with us