Building a Culture of Documentation

Every engineering organization I have led has reached the same uncomfortable moment. A senior engineer leaves, a critical incident hits at two in the morning,...

Originally published onanselmfowel.com

Every engineering organization I have led has reached the same uncomfortable moment. A senior engineer leaves, a critical incident hits at two in the morning, or a regulator asks a question about a decision made three years ago, and the answer lives only inside someone's head. In a regulated fintech business, that gap is not a minor inconvenience. It is an operational risk, a compliance exposure, and a tax on every future decision. The fix is not a wiki nobody reads. The fix is culture.

Building a Culture of Documentation
Building a Culture of Documentation

I want to be precise about what I mean by a culture of documentation, because the phrase gets abused. I do not mean producing more documents. I mean building an organization where writing things down is a default behavior, where the cost of doing it is low, and where the value compounds over time. That is a leadership problem far more than a tooling problem, and I have learned most of what I know about it by getting it wrong first.

Why Documentation Is a Leadership Problem, Not a Tooling Problem

For years I believed documentation was a discipline issue. If engineers were just more conscientious, the thinking went, the docs would exist. So we bought better tools, ran campaigns, and added a documentation step to the definition of done. None of it stuck, because none of it addressed the actual incentives. People do what is rewarded and skip what is punished, and in most engineering cultures, shipping is rewarded and writing is invisible.

The shift happened when I started treating documentation as a function of how we organize work rather than a personal virtue. Engineers are not lazy about writing. They are responding rationally to an environment where the document they spend two hours on is never read, never referenced in a review, and never credited in a promotion case. If you want different behavior, you change the environment, and that is squarely the job of leadership.

This reframing matters because it tells you where to spend your energy. You do not need to lecture people about the importance of knowledge sharing. You need to remove friction, create reasons to read what gets written, and make the act of documenting feel like part of the work rather than an interruption to it. Everything that follows is downstream of that one decision.

The Real Cost of Undocumented Knowledge in Regulated Fintech

In a payments or lending business, undocumented knowledge is not a soft cost. It shows up in concrete, expensive ways. When a payment reconciliation job behaves strangely and the only person who understands the rounding logic is on holiday, the on-call engineer is reverse-engineering production behavior under time pressure with customer money in the balance. That is precisely the situation that turns a small anomaly into a reportable incident.

There is also the regulatory dimension. Auditors and supervisors increasingly expect to see not just that controls exist but that decisions were reasoned and recorded. When a regulator asks why a particular fraud threshold was set, the answer cannot be that someone tuned it by feel in 2022 and has since left the company. We need a defensible record of the rationale, the data considered, and the approvals obtained. Documentation is, in this context, part of the control environment itself.

The most expensive code is not the code that breaks. It is the code that works perfectly and that no one understands well enough to change.

I have watched teams slow to a crawl not because their systems were poorly built but because the knowledge required to modify them safely had evaporated. The original authors moved on, the design decisions were never captured, and the surviving engineers treated the code as load-bearing infrastructure they were afraid to touch. That fear is the true cost, and it accumulates silently until you try to make a change and discover you cannot move.

Designing Documentation as a Byproduct of Work

The single most effective change I have made is to stop asking for documentation as a separate deliverable and start capturing it as a byproduct of work that has to happen anyway. People resist writing a document. They do not resist writing a thorough pull request description, a decision record attached to a design they were going to present regardless, or a clear ticket that explains the why. The trick is to attach the durable record to the activity that already exists.

Concretely, this means a few things in how we run engineering:

  • Design discussions produce a short written decision record before code is merged, not a meeting that evaporates into nothing.
  • Pull requests require a description of intent and tradeoffs, and reviewers are expected to push back when that section is thin.
  • Incident reviews generate a written timeline and root cause that lands in a searchable place, not a private channel thread.
  • Runbooks are written when a procedure is first performed manually, while the steps are still fresh, rather than recalled months later.

None of these are heavyweight. Each one is a small amount of writing attached to a moment when the author already has the context loaded in their mind. The cost of capturing knowledge at that moment is a fraction of the cost of reconstructing it later, and the quality is far higher because nobody is guessing about what they were thinking.

Lightweight Decision Records Over Heavy Specifications

I am deeply skeptical of the heavyweight specification document. The two-hundred-page design spec is a relic of a world where software changed slowly. In practice these documents are obsolete the moment they are approved, they are rarely read in full, and the effort of maintaining them is so high that nobody bothers. They give the appearance of rigor without the substance.

What works instead is the architecture decision record: a short, dated note that captures a single decision, the context that forced it, the options considered, and the consequences accepted. A good decision record fits on one screen. It does not describe the system in exhaustive detail. It captures the reasoning at a fork in the road, which is exactly the thing that is impossible to reconstruct later and exactly the thing future engineers most need to understand.

The reason this format succeeds is that it respects the reader and the writer. The writer can produce it in fifteen minutes because it asks only for the reasoning they already did. The reader can consume it quickly and trust that it reflects a real moment in time rather than an aspirational design. When someone asks why our ledger uses a particular consistency model, I can point them to a record written the week we decided it, and they can see not just the answer but the constraints that produced it.

Writing for the Engineer at Two in the Morning

The most honest test of documentation quality is whether it helps a tired, stressed engineer at two in the morning. That person does not have time to read a narrative essay about your system philosophy. They need to know what is broken, what normal looks like, what levers are safe to pull, and who to escalate to. Operational documentation written for that reader looks very different from documentation written to impress a reviewer.

I push teams to write runbooks in the imperative, with concrete commands and concrete expected outputs, and to assume the reader knows nothing about the system's history. The goal is to let someone with general competence but no specific context recover a service safely. If a runbook requires you to already understand the system to follow it, it has failed at the one job that matters most. We test this directly by having someone unfamiliar with a service follow its runbook during a game day.

This reader-centric framing also clarifies what does not need to be written. Not every internal detail deserves a document. The questions worth answering in writing are the ones that recur, the ones that are dangerous to get wrong, and the ones where the answer is non-obvious. Writing everything is as useless as writing nothing, because volume buries the few documents that genuinely matter under a pile of noise.

Making Documentation Discoverable, Or It Does Not Exist

A document that cannot be found does not exist. I have seen organizations with genuinely good documentation that nobody used because it was scattered across three wikis, a shared drive, an old ticketing system, and a folder of screenshots in a chat tool. The knowledge was technically present and practically unavailable, which is worse than absent because people stop trusting that searching is worthwhile.

Discoverability is an architectural concern that deserves real attention. We standardized on a single source of truth for each kind of knowledge, with predictable locations and consistent naming, so that an engineer can guess where something lives and usually be right. Search has to work, links have to be stable, and there has to be an obvious entry point for a service that leads to its design records, its runbooks, and its dependencies. When the structure is predictable, people stop asking colleagues and start finding answers themselves.

I also insist on pruning. Stale documentation is actively harmful because it teaches people that the documentation lies, and once they believe that, they stop reading all of it. We treat a wrong runbook as a defect to be fixed or deleted, not a minor cosmetic issue. A smaller, trusted body of documentation beats a vast, suspect one every time, and maintaining that trust is an ongoing act of curation rather than a one-time effort.

Leaders Set the Standard by Writing First

Culture is set by what leaders do, not what they say, and documentation culture is no exception. If I ask engineers to write decision records while I make architectural calls in hallway conversations that never get captured, the message is clear: documentation is for junior people, not for the people whose decisions matter most. I have to write, visibly and consistently, and I have to reference what others have written.

One of the most powerful signals is to read documentation in public and act on it. When I cite a colleague's design record in a review, ask a clarifying question against a runbook, or point a stakeholder to an existing decision rather than re-deciding it, I demonstrate that writing things down produces leverage. People write for an audience, and the moment they learn there is a real audience, the quality and frequency of their writing improves on its own.

I also try to make good writing a recognized form of engineering excellence. In promotion discussions and performance reviews, the engineer who left behind clear records that allowed others to move faster is doing high-leverage work, and I name it as such. The engineer who builds something brilliant that only they can maintain has created a liability, however impressive the artifact. Rewarding the right thing is the lever that turns a campaign into a durable norm.

Documentation in the Age of AI Assistants

The arrival of capable coding assistants has changed the calculus, though not in the way some people assume. The naive hope is that we can simply generate documentation on demand and stop worrying about writing it ourselves. In practice, a model can describe what code does, but it cannot reliably tell you why a decision was made, what alternatives were rejected, or what constraint from a regulator shaped a particular design. That intent lives in people, and it has to be captured by people.

What has genuinely improved is the cost of producing the mechanical parts. Summarizing a change, drafting the first version of a runbook from a sequence of steps, or generating a clear description of an interface are tasks that an assistant can accelerate, leaving engineers to focus on the reasoning that only they possess. I encourage teams to use these tools to lower the friction of the routine writing, then spend the saved effort on the parts that carry real judgment. The human contribution shifts toward intent and away from transcription.

There is a second-order effect worth noting. Well-structured documentation is now also an input to the tools our engineers use, because assistants reason better about a codebase that has clear decision records and accurate operational notes. The investment in writing things down clearly pays off twice: once for the human reader and once for the systems that increasingly help us work. That reinforces, rather than replaces, the discipline of capturing knowledge.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

A culture of documentation is not built by mandate, and it is not bought with a tool. It is built by leaders who treat written knowledge as part of the work, who design the organization so that capturing it is cheap and reading it is valuable, and who model the behavior themselves rather than delegating it downward. In a regulated fintech business, where the cost of forgotten reasoning is measured in incidents, audits, and frozen systems, this is not an optional nicety. It is core operational hygiene. The teams that internalize it move faster, recover from failures more calmly, and answer hard questions with confidence, because the knowledge they rely on does not walk out the door when a single person does.

Chat with us