Why Burnout Is a System Problem

For most of my career I treated burnout the way I treated a flaky integration: as a symptom of someone, somewhere, not being careful enough. If an engineer was...

Originally published onanselmfowel.com

For most of my career I treated burnout the way I treated a flaky integration: as a symptom of someone, somewhere, not being careful enough. If an engineer was exhausted, surely they had taken on too much, failed to set boundaries, or skipped the gym for too many weeks. The fix, I assumed, was a conversation about self-care and a few days of paid time off. It took watching several good engineers leave roles they had once loved before I understood how wrong that framing was.

Why Burnout Is a System Problem
Why Burnout Is a System Problem

Burnout is not a personal failing that occasionally appears inside an otherwise healthy system. In the teams I have led, it is the predictable output of how the system is designed. When the same patterns produce the same exhaustion across different people, the people are not the variable. The system is. This post is my attempt to lay out what I have learned about treating burnout as an engineering problem with engineering causes, because that reframing is the only one that has ever led me to a durable fix.

Why the Individual Framing Fails

The dominant story about burnout puts the cause and the cure inside the individual. The engineer is tired, so the engineer needs rest, resilience training, or better time management. This framing is seductive because it is cheap. It costs a manager nothing structural to suggest a meditation app. But it consistently fails, and the failure mode is instructive: the person rests, returns, and burns out again on the same schedule, because nothing about the environment that produced the exhaustion has changed.

I have watched this play out enough times to recognize the shape of it. A strong contributor goes on leave, comes back recharged, and within two months is right back where they started. If rest were the cure, the rest would hold. The fact that it does not tells me the cause lives outside the person.

If you replace a burned-out person with a fresh one and the fresh one burns out on the same timeline, you have not diagnosed a people problem. You have characterized a workload.

Chronic Overload Is a Design Choice

The most common structural cause I see is simply that the work does not fit the people. Not for a sprint, not for a quarter-end crunch, but as a steady state. A team of six is staffed and planned as though it were a team of nine. Every planning meeting quietly assumes that everyone will operate at peak output every week, with no slack for illness, interrupts, or the ordinary friction of building software in a regulated environment.

This is a design choice, even when nobody consciously decides it. When a roadmap is built on the assumption of one hundred percent utilization, you have designed a system with no margin. Any real system needs slack to absorb variance, and software teams are no exception. The queueing theory is not subtle here: as utilization approaches its ceiling, wait times and stress climb toward infinity. A highway at one hundred percent capacity is not efficient, it is a traffic jam. Teams behave the same way.

In fintech the problem compounds, because the work is not only feature delivery. There is audit preparation, incident response, regulatory change, and the constant hum of compliance obligations that never appear on a roadmap but consume real hours. If those hours are unplanned, they come out of the same finite budget as everything else, and the budget was already fully spent before the unplanned work arrived.

Ambiguity Is a Hidden Tax

Overload is the obvious culprit, but it is rarely the only one. A quieter cause is chronic ambiguity: unclear ownership, shifting priorities, and goals that change faster than work can be completed. An engineer who finishes a two-week effort only to learn it was deprioritized mid-flight has not just lost time. They have lost the sense that their effort connects to an outcome, and that connection is one of the load-bearing beams of motivation.

I think of ambiguity as a tax on every action. When ownership is unclear, every decision requires a round of negotiation about who decides. When priorities shift weekly, every plan carries the unspoken expectation that it will be discarded. The cognitive overhead of constantly re-deriving context is enormous, and unlike feature work it produces nothing visible. People feel that they are running hard and getting nowhere, one of the most reliable precursors to burnout I know.

The remedy is not more process for its own sake. It is clarity about a few things: who owns this, what are we optimizing for this quarter, and what will we not do. The teams I have led that handled stress best were rarely the ones with the lightest workload. They were the ones who knew why they were doing what they were doing.

The On-Call Asymmetry

On-call deserves its own discussion because it is where so many burnout trajectories begin, and because its cost is systematically underestimated. The visible cost of on-call is the pages. The invisible cost, far larger, is the tax it places on time that is nominally free. An engineer carrying the pager cannot fully relax, cannot have a glass of wine with dinner, cannot take their child to a weekend event without one eye on their phone. That weekend is not rest even if no page arrives.

When the rotation is thin, this asymmetry becomes corrosive. A four-person rotation means one week in four is compromised, every month, indefinitely. Add a noisy alerting setup that pages for non-actionable events at three in the morning, and you have engineered a machine for grinding people down. The engineers are not weak for finding it unsustainable. The rotation is unsustainable, and no amount of grit changes that arithmetic.

  • Rotations thinner than six to eight people compromise rest at a rate most people cannot sustain for years.
  • Every non-actionable page is a withdrawal from a finite trust account, and that account does not refill on its own.
  • Sleep interrupted by alerts is not recoverable by sleeping in the next day; the damage is to the night, not the total hours.
  • On-call load that is never reviewed against actual incident volume will silently drift toward unbearable.

Recognition, Agency, and the Sense of Progress

Two structural factors that rarely make it into burnout conversations are the absence of recognition and the absence of agency. Recognition here does not mean praise theater. It means that the work people do is seen, understood, and connected to something that matters. When an engineer spends a month hardening a payment reconciliation path that prevents a class of errors no customer will ever notice, and that work is invisible to everyone above them, the message they receive is that invisible work does not count. They will stop doing it, or they will leave.

Agency is the other half. People can tolerate a remarkable amount of hard work when they feel they have some control over how it is done. They tolerate very little when they feel like a ticket-processing machine, executing decisions made elsewhere with no input into the how or the why. The combination of high demand and low control is, in the occupational health literature, one of the most reliable predictors of strain. I have seen it in my own teams: the most exhausted people were not always the busiest, they were the ones who felt the work was happening to them.

As a leader, the levers here are within reach. I can make sure the unglamorous, system-strengthening work is named and valued in reviews, and I can push real decisions down to the people doing the work rather than hoarding them. Neither requires budget. Both require me to pay attention to where effort goes and who gets to shape it.

Leadership Behaviors That Leak Downward

I have to be honest about my own contribution, because leadership behavior propagates through a team with surprising fidelity. If I send messages at midnight, some fraction of my team will conclude that midnight availability is expected, no matter what I say otherwise. If I celebrate the heroic weekend rescue more loudly than the boring quarter where nothing broke, I am training people to value heroics over sustainability, and heroics are burnout in a flattering costume.

The hardest version of this is that exhaustion at the top becomes exhaustion everywhere. A burned-out manager makes worse decisions, communicates less clearly, and models a relationship with work that the team absorbs. I cannot ask a team to maintain sustainable pace from a position of visible depletion. The system includes me, and my habits are inputs to it whether I intend them to be or not.

A team rarely rises above the relationship its leader has with rest. If I treat my own recovery as optional, I have made it optional for everyone watching.

Instrumenting the System for Early Signals

Because burnout is a system property, it can be measured, at least indirectly. I do not mean surveillance, which backfires badly. I mean watching the same kind of leading indicators I would watch for any production system whose health I cared about. The signals are usually there well before someone hands in a resignation; the problem is that nobody is looking at them as a connected story.

The patterns I watch for are simple. Commit and review activity creeping into nights and weekends. Pull requests sitting longer because the reviewers are underwater. Planned-to-actual capacity ratios that never come back into balance. A rise in small mistakes from people who normally do not make them, often the first sign of accumulated fatigue. None is conclusive alone, but together they form a picture, and the picture tends to precede the crisis by months.

The discipline is to treat these as system metrics rather than performance metrics. If review latency is climbing, the answer is not to tell people to review faster. It is to ask why the system is producing more work than it can absorb. Pointing the metric at the individual recreates exactly the framing error this whole post is arguing against.

Redesigning for Sustainability

The good news about a system problem is that systems can be redesigned. The interventions that have worked for me are unglamorous and structural rather than motivational. I plan for something well below full utilization, leaving genuine slack so that variance does not immediately translate into overtime. I fund the unsexy work, the alert tuning and the test infrastructure, because it is what keeps the on-call rotation from becoming a grinder.

I have learned to protect rotations by refusing to let them get too thin, even when that means saying no to scope. I make priorities explicit and stable enough that a two-week effort can survive to completion. And I treat my own working hours as a signal others read, because they do. None of this is a perk or a wellness initiative. It is the ordinary work of designing a system that can run for years without consuming the people in it.

The reframe matters because it changes what you do on Monday morning. If burnout is a personal failing, your tools are pep talks and time off. If it is a system problem, your tools are staffing, scope, alerting hygiene, and clarity of ownership. The second set is the only one I have ever seen produce a team that stays healthy under sustained load.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

I no longer ask why a particular engineer is burning out. I ask what about the system makes burnout the predictable result, because that is the question with answers I can act on. The shift from blaming people to examining design is the single most useful change I have made as a leader, and it is the one I would urge most on anyone running a team under real pressure. Build the system so that staying is sustainable, and the people problem you thought you had will turn out to have been an architecture problem all along, the kind you actually know how to fix.

Chat with us