When I started in fintech, the word itself still carried a faint whiff of novelty. We were the upstarts, the people building software around the edges of banks that had run on mainframes since before I was born. A decade later, the novelty is gone and the responsibility has hardened. I have shipped payment rails, watched a reconciliation bug quietly misroute money for six hours, sat across from regulators, and rebuilt a fraud system twice. What follows is not a victory lap. It is the set of lessons that survived contact with reality.
I am writing this for the engineer who is three years in and starting to suspect that the hard parts of fintech are not the algorithms. You are right. The hard parts are the ones nobody puts on a slide: correctness under partial failure, the cost of trust, and the discipline to move slowly in the places that matter while moving quickly everywhere else.
Money Is Not a Number
The first thing fintech teaches you, usually painfully, is that money is not a number in a database. It is a claim, a promise, and a legal obligation, all of which have state that lives outside your system. A balance is the sum of events you have agreed to honor, not a field you are free to overwrite. I have seen more incidents caused by treating a balance as mutable state than by any clever attack.
This is why double-entry accounting, a four-hundred-year-old idea, still outperforms most of the ledger designs I have reviewed. Every movement has two sides; the books must balance or something is wrong. When we adopted an append-only ledger where balances were derived rather than stored, our entire class of "the number is wrong and we cannot explain why" incidents disappeared. We traded a little query performance for the ability to always answer the question that matters most: how did this account get to this state?
The corollary is that you must respect representation. Floating point has no place near currency. We standardized on integer minor units everywhere, with currency carried alongside the amount as a first-class value rather than an afterthought. It sounds trivial until you debug a rounding discrepancy that only appears across a fee calculation in a third currency.
Idempotency Is a Feature, Not a Trick
Networks fail in the middle of operations. A client sends a payment request, the connection drops, and now neither side knows whether the money moved. The naive answer is to retry. The naive answer doubles people's payments and ends careers.
Every state-changing endpoint we own takes an idempotency key, and the server is responsible for guaranteeing that the same key produces the same result exactly once, no matter how many times it arrives. This is not a nice-to-have. It is the load-bearing wall of a payments system. We learned to treat the idempotency layer as a product surface with its own tests, its own retention policy, and its own dashboards, because callers depend on it more than they depend on any business feature.
The question is never "did the request succeed?" It is "can the caller find out what happened without making it happen again?" If the answer is no, you have built a system that punishes the honest behavior of retrying.
Failure Is the Normal Case
Early on I designed for the happy path and bolted on error handling. That is backwards. In fintech, the interesting states are the in-between ones: pending, reversed, partially settled, disputed, clawed back. A transaction that simply succeeds is the least of your concerns. The money that is stuck in an ambiguous state at two in the morning is what wakes you up.
So we started modeling lifecycles explicitly. Every payment is a state machine with named transitions, and we forbid implicit transitions entirely. If a payment can go from authorized to captured, there is a function for it, with preconditions, and there is no other way to get there. When a counterparty's webhook arrives twice, or out of order, or three days late, the state machine either accepts it idempotently or rejects it loudly. There is no quiet corruption.
The mindset shift is to stop asking "what could go wrong" and start assuming everything will, then deciding in advance what each failure means for the money. A timeout is not an error to log and forget. It is a fork in the ledger that must be resolved by reconciliation, by a reversal, or by a human, but never by silence.
Reconciliation Is Where Truth Lives
Your system thinks it moved money. The bank thinks something slightly different. The card network has a third opinion, and it arrives in a file at four in the morning in a format designed in 1987. The only way to know what actually happened is reconciliation, and the teams that take it seriously sleep better than the teams that treat it as a back-office afterthought.
We built reconciliation as a first-class pipeline that runs continuously, not a monthly scramble. Every internal event is matched against external statements, and anything unmatched past a threshold raises an alert before a customer ever notices. The earliest version of this caught a settlement file we had been silently double-counting for a week. The numbers were small, but the trust implications were not.
- Reconcile against the source of record, not against another copy of your own data.
- Make breaks visible and assignable, not buried in a report nobody reads.
- Measure the age of the oldest unreconciled item; that number is a health metric for the whole company.
- Automate the matching, but keep a human in the loop for the exceptions, because exceptions are where fraud and bugs hide.
Compliance Is an Engineering Discipline
I used to think of compliance as a tax the lawyers imposed on the fun work. That attitude is expensive. Regulation in financial services is, for the most part, a compressed record of ways that companies have previously harmed customers. Read that way, it is closer to a postmortem library than to bureaucracy.
The teams that struggle are the ones that treat audit and control as something you produce at the end, by hand, under duress. The teams that thrive build evidence generation into the system. Access to production data is logged because the logging is the control, not because someone will ask for it later. Change management is enforced in the pipeline. When an auditor asks who approved a given deployment, the answer is a query, not an archaeology project.
My rule of thumb is simple: if a control depends on a person remembering to do something, it is not a control, it is a wish. Encode the requirement in the system and you turn a recurring source of anxiety into a property you can prove on demand.
Security Is a Product Property
In a payments company, security is not a feature you add; it is a constraint that shapes every decision, the way load-bearing walls shape a building. The threat model is genuinely adversarial. There are people whose full-time job is to take money from your customers through your software, and they are creative, patient, and well funded.
The practical lesson is that defense in depth beats cleverness. We assume any single layer can fail, so we make sure no single failure is catastrophic. Sensitive credentials are scoped tightly and rotated automatically. The blast radius of any one compromised service is bounded by design. We invested early in the unglamorous work of secrets management and least-privilege access, and it paid for itself the first time an exposed key turned into a non-event because that key could do almost nothing.
The most dangerous phrase in a security review is "no one would ever do that." Someone will, and they will document it on a forum, and then everyone will.
Move Fast Where It Is Safe, Slow Where It Is Not
There is a tired debate about whether fintech teams should move fast or move carefully. The answer is both, applied to different parts of the system. The mistake is to apply one tempo uniformly. A marketing page and a settlement engine do not deserve the same change process, and treating them the same either grinds your product team to a halt or puts your customers' money at risk.
We drew an explicit map of the codebase by consequence of failure. The core ledger, the movement of funds, the authentication boundary: these change slowly, with review, with extensive testing, with rollback plans rehearsed in advance. The experience layer around them moves quickly, ships continuously, and is free to fail in small, recoverable ways. The art of fintech engineering leadership is keeping that boundary crisp and refusing to let the slow rigor bleed into places that do not need it, or the fast experimentation creep into places that cannot tolerate it.
This is also a hiring and culture decision. Engineers need to internalize which side of the line they are working on today, and feel equally comfortable shipping ten times a day on one side and spending two weeks on a single migration on the other.
Incidents Are the Real Curriculum
No course taught me as much as the incidents did. The six-hour misrouting bug I mentioned earlier reshaped how I think about deployments, observability, and the humility required to operate systems that touch money. We caught it through reconciliation, not through monitoring, which told us our monitoring was looking at the wrong layer.
What separated the organizations I respect from the ones I do not was their relationship with failure. The good ones ran blameless postmortems that produced concrete, owned action items and actually closed them. The bad ones produced documents that assigned fault and changed nothing, which guaranteed the same incident would return wearing a different costume. A postmortem that ends with "be more careful" is a postmortem that has failed.
The discipline I carry from this is to treat every incident as a paid lesson, expensive enough that wasting it is the real failure. We track not just whether action items exist but whether they ship, and we revisit old incidents to ask whether the fix held. Reliability is not a state you reach; it is a practice you maintain.
Trust Is the Actual Product
After ten years I am convinced that the product we sell is not payments, accounts, or dashboards. It is trust, denominated in correctness and availability. People hand us their money and the data that describes their financial lives because they believe we will not lose it, leak it, or misplace it. Every technical decision is, underneath, a decision about whether that belief is justified.
This reframes a lot of arguments. A feature that ships a day late costs a day. A feature that erodes trust can cost a company. That asymmetry is why fintech engineering rewards a particular temperament: ambitious about what to build, conservative about how the money moves, and allergic to the phrase "it probably works." Probably is not a number you can put in a ledger.

Conclusion
If I compressed the decade into a single sentence, it would be this: in fintech, the engineering and the responsibility are the same job. The state machines, the idempotency keys, the reconciliation pipelines, and the audit trails are not separate from the duty of care; they are how that duty is expressed in code. The teams that understand this build systems that are boring in the best possible way, where the money moves exactly as promised and the surprises stay small. That is the standard I hold myself to now, and the one I would offer to anyone three years in, still discovering that the hard parts were never the algorithms.
