The Habits of Calm Engineering Teams

I have managed engineering teams through payment outages on the busiest shopping day of the year, through regulatory audits that arrived with two weeks notice,...

Originally published onanselmfowel.com

I have managed engineering teams through payment outages on the busiest shopping day of the year, through regulatory audits that arrived with two weeks notice, and through the quiet, grinding pressure of building reconciliation systems that simply cannot be wrong. Over those years I have come to believe that the single most reliable predictor of a team's long-term output is not its raw talent, its tooling, or even its architecture. It is the team's baseline emotional temperature. Calm teams ship more, break less, and stay together longer.

The Habits of Calm Engineering Teams
The Habits of Calm Engineering Teams

Calm is not the absence of urgency. In regulated fintech there is plenty of urgency, and there are nights when something genuinely is on fire. Calm is the practiced capacity to meet urgency without panic, to absorb bad news without flinching, and to keep thinking clearly when the stakes are high. That capacity is built deliberately, through habits, long before the pressure arrives. Here are the habits I have seen distinguish the calm teams from the frantic ones.

They Treat Incidents as Information, Not Indictments

The fastest way to make a team anxious is to make every incident a referendum on someone's competence. When a failed deployment or a missed reconciliation triggers a search for the guilty party, people learn to hide problems, delay disclosure, and shade their language. That is precisely the opposite of what you want when real money is moving and a small discrepancy today becomes a regulatory finding next quarter.

Calm teams run blameless postmortems, but they go further than the ritual. They genuinely believe that an incident is a gift of information about how the system actually behaves under load, not how the architecture diagram says it should behave. When a settlement job double-posted because of a retry storm, the question is never who pressed the button. The question is why the system made it so easy to press the wrong button, and why nothing caught it for forty minutes.

I have watched this shift transform a team's reporting behavior within a single quarter. Once engineers trusted that raising their hand early would be rewarded rather than punished, the average time-to-disclosure on issues dropped dramatically. We caught more problems while they were small, which is the entire game in payments.

They Protect Cognitive Headroom Deliberately

Panic is partly a resource problem. A mind that is already running at full capacity has nothing left to spend when something unexpected arrives, so the unexpected feels like a crisis. Calm teams treat cognitive headroom as a managed resource, the same way they treat database connections or rate limits. They do not run their people at one hundred percent utilization and then act surprised when a minor incident causes a meltdown.

In practice this means accepting that a sustainable team operates with visible slack. Sprints are not packed to the last hour. On-call rotations are sized so that the person carrying the pager is not also expected to deliver a feature that week. There is room in the schedule for the unplanned, because in any system that touches real-world money, the unplanned is not an edge case. It is a recurring line item.

A team running at full utilization has no capacity to think. It can only react. And reaction, in a regulated environment, is how small mistakes become reportable events.

This is a hard sell to finance-minded stakeholders who see slack as waste. The reframe I use is that you are not paying for idle time, you are paying for response capacity. An ambulance that is busy ninety-five percent of the time is a public health failure, not an efficiency triumph. The same logic applies to the people who keep your payment rails running.

They Make the On-Call Experience Humane

Nothing erodes a team's calm faster than a brutal on-call rotation. If being on the pager means a week of fragmented sleep, alert fatigue, and dread, then your most experienced people will quietly start looking for the exit, and the ones who remain will be too exhausted to think clearly during the incidents that matter. On-call humaneness is not a perk. It is a reliability investment.

The teams I trust most are ruthless about alert quality. Every page that fires must be actionable and must genuinely require a human in the next few minutes. Anything that does not meet that bar gets demoted to a dashboard or a ticket. I have seen rotations where seventy percent of overnight pages were noise, and the cost was not just lost sleep. It was that the engineer had learned to swipe alerts away half-asleep, and eventually swiped away the one that mattered.

  • Every alert links directly to a runbook with concrete first steps, so the responder is never starting from a blank page at three in the morning.
  • Pages that cannot be acted upon immediately are reclassified the same week, not left to accumulate.
  • Anyone who takes a rough overnight incident gets the next day off without negotiation or guilt.
  • The on-call handoff includes a written summary of anything still simmering, so the incoming engineer inherits context, not surprises.

They Default to Written, Asynchronous Thinking

Frantic teams live in meetings and instant messages. Decisions get made verbally, context evaporates, and the same arguments resurface every few weeks because nobody can remember what was concluded or why. The constant interruption of synchronous communication also fragments the deep focus that hard engineering problems require. You cannot reason carefully about a distributed transaction protocol in the gaps between four standups.

Calm teams write things down. Design documents, decision records, and incident reviews are not bureaucratic overhead, they are the institutional memory that lets the team stay calm when a question resurfaces eight months later. When someone asks why we chose eventual consistency for a particular ledger view, the answer is a link, not a frantic reconstruction from people's faulty recollections.

Writing also slows thinking down in a productive way. It is much harder to make a sloppy argument in prose than in a hallway conversation, because the gaps in your reasoning become visible on the page. A team that writes its proposals catches its own bad ideas before they reach production, which means fewer of the late-night reversals that drain a team's composure.

They Separate Decision Pressure From Execution Pressure

A subtle source of anxiety is conflating the urgency of a decision with the urgency of its execution. A regulatory deadline might be genuinely fixed, but that does not mean every architectural choice supporting it must be made in a hurry. Calm teams are precise about which clock they are actually racing, and they refuse to let an artificial sense of haste contaminate decisions that deserve more thought.

When a payment scheme announces a mandate change with a hard compliance date, the temptation is to treat everything connected to it as equally urgent. The disciplined response is to ask what genuinely must be decided now, what can be reversed cheaply later, and what is merely loud. Most architectural decisions are more reversible than they feel in the moment, and treating them as one-way doors when they are actually two-way doors manufactures stress for no benefit.

I encourage my teams to name the deadline explicitly and then ask what the latest responsible moment is to make each decision. More often than not, the real constraint is days or weeks away, not the same afternoon, and simply saying that out loud lowers the temperature in the room enough for people to think properly.

They Normalize Saying I Do Not Know Yet

In high-stakes environments there is enormous pressure to project certainty. An engineer asked about the blast radius of a change feels obligated to give a confident answer even when the honest answer is that they are not sure. This false confidence is corrosive, because it means decisions get made on fabricated certainty, and the team is repeatedly blindsided by consequences that someone could have flagged if the culture had allowed honesty.

Calm teams make I do not know yet a complete and respectable sentence. They treat uncertainty as a thing to be investigated rather than a personal failing to be concealed. When an engineer says they are not sure whether a migration is safe under concurrent load, that is not weakness. That is the most valuable signal in the room, and a calm team responds by funding a spike to find out rather than by pressuring the person to guess.

This habit compounds. Once it is safe to admit uncertainty, people stop wasting energy maintaining a facade of omniscience and start spending that energy on actually reducing the uncertainty. The team's collective map of what it does and does not understand becomes accurate, and an accurate map is the foundation of every calm decision.

They Design Systems That Fail Gently

Team calm and system design are deeply intertwined. A brittle system that fails catastrophically and opaquely will keep a team in a permanent state of low-grade dread, because everyone knows the next failure could be the one that loses money or trust. A system designed to fail gently, in contrast, gives a team room to breathe even when something breaks.

Concretely, calm teams invest in the unglamorous properties that make failures survivable. Idempotency so that a retry cannot double-charge a customer. Circuit breakers so that a struggling downstream dependency degrades service rather than collapsing it. Reconciliation processes that catch discrepancies automatically rather than relying on a customer complaint to surface a missing payment. These properties are not where the excitement lives, but they are where the calm lives.

The deepest version of this principle is that the system should make the safe path the easy path. If deploying carefully requires heroics while deploying recklessly is one click, your team will eventually be reckless under pressure, no matter how disciplined the individuals are. Calm is partly a property of an environment that gently steers tired people toward safe behavior when their own judgment is depleted.

They Have Leaders Who Absorb Pressure Rather Than Transmit It

Anxiety flows downhill through an organization with remarkable efficiency. When a board is nervous and a CTO transmits that nervousness directly to engineering, the team inherits a level of fear wildly disproportionate to anything they can actually act on. One of the core jobs of a leader in a high-stakes domain is to act as a filter and a buffer, absorbing the raw pressure from above and passing down only the part that is genuine, actionable, and accompanied by enough context to be useful.

This does not mean hiding reality from the team. Treating engineers like adults means being honest about genuine business pressure and real risk. The distinction is between transmitting information and transmitting emotion. The team needs to know that a key client is unhappy with our uptime. They do not need to absorb the panic of the third escalation call I sat through this morning. I take that emotional load so they can focus on the engineering problem with a clear head.

The leaders who do this well are usually the ones who are visibly calm during incidents. Their steadiness during the worst moments is not performance, it is the most important thing they contribute. A leader who stays measured while a payment outage unfolds gives the responding engineers permission to stay measured too, and measured engineers resolve incidents faster than frightened ones.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

None of these habits is exotic, and none requires a particular technology stack or a generous budget. They require something harder, which is sustained intention over time, especially during the periods when everything is going well and the temptation to cut the slack and pack the schedule is strongest. Calm is built in the quiet weeks so that it is available in the loud ones. The teams that invest in it discover that calmness is not the opposite of high performance. In any domain where the cost of a careless mistake is measured in real money and real trust, calmness is the precondition for it.

Chat with us