Three years ago I inherited a payments platform that had been split into 47 microservices by a team of nine engineers. Nine people, 47 services. A single card authorization touched eleven of them, and each hop added a network call, a retry policy, a serialization boundary, and one more place for a 3am page to originate. The founders were proud of it. They had read the same conference talks everyone reads, and they had built the architecture those talks described. What they had actually built was a distributed monolith with worse latency and no transactions.
I am not against microservices. I have run platforms where they were exactly right. But I have come to believe that most teams reach for them years too early, for reasons that have very little to do with the problems microservices actually solve. This is a post about that mistake, what it costs, and how to tell whether you are about to make it.
The resume-driven default
Somewhere around 2015, "we use microservices" stopped being an architecture decision and became a hiring signal. Candidates expect it. Engineers want it on their CV. A senior developer who spent two years wiring up service meshes is more marketable than one who kept a well-factored monolith healthy, even though the second job is harder and more valuable. So the incentives push toward fragmentation before there is any technical reason to fragment.
I have sat in the architecture review where a lead proposed breaking out a "notifications service" for a product with 4,000 users. The real driver was not scale. It was that notifications felt like a clean bounded context and splitting it out looked like good engineering. Nobody says the resume part out loud, which is why it slips through. It arrives dressed as separation of concerns, or future-proofing, or "we'll want to scale this independently one day." That day, for most companies, never comes.
What you actually pay for the split
The bill for premature microservices does not arrive as one big invoice. It arrives as a hundred small taxes you stop noticing because you assume they are just the cost of software. They are not. They are the cost of a decision.
- Every in-process method call that used to be a compiler-checked function is now a network call that can time out, return a 503, or silently retry and double-charge someone.
- A schema change that was a single migration is now a coordinated deploy across three teams with a backward-compatibility window.
- You lose database transactions across service boundaries, so you rebuild them by hand as sagas and compensating actions, badly, and then you spend Q3 chasing the states they leave stranded.
- Local reproduction of a bug now requires standing up eight services, or an elaborate mock harness that drifts from reality within a month.
- Your observability bill triples because you cannot understand a single request without distributed tracing.
None of these are hypothetical. At that 47-service shop, our median build-and-deploy for a one-line pricing fix was 40 minutes across the affected services, and a full end-to-end test environment took the better part of an afternoon to bring up. We had traded the discomfort of a large codebase for the far worse discomfort of a large network.
The transaction you gave away
The loss I feel most sharply in fintech is the database transaction. In a monolith, moving money and recording the ledger entry happen in one atomic unit. Either both commit or neither does. It is boring, it is correct, and the database has spent forty years making it fast.
Split the ledger and the wallet into two services and that guarantee is gone. Now you are hand-rolling consistency. Here is the kind of thing you find yourself writing, and the kind of thing that keeps a regulated business awake:
// Monolith: one transaction, provably correct
using var tx = await _db.BeginTransactionAsync();
await _wallet.DebitAsync(accountId, amount);
await _ledger.RecordAsync(new Entry(accountId, amount, "debit"));
await tx.CommitAsync();
// Two services: no shared transaction, so you improvise
var reservation = await _walletClient.ReserveAsync(accountId, amount);
try
{
await _ledgerClient.RecordAsync(new Entry(accountId, amount, "debit"));
await _walletClient.ConfirmAsync(reservation.Id);
}
catch (Exception)
{
// If THIS call fails too, you now have a reservation
// stranded in an unknown state. Have fun reconciling.
await _walletClient.ReleaseAsync(reservation.Id);
throw;
}
The second version is not wrong, exactly. It is what you must do once you have made the split. But look at what you signed up for: a reservation that can strand, a release that can itself fail, and a reconciliation job that has to reason about every partial state. In a payments context, a stranded reservation is not an abstraction. It is a customer's money that is neither spent nor available, and eventually it is a support ticket, and occasionally it is a regulator asking why.
Conway's law cuts both ways
The strongest honest case for microservices is organizational, not technical. When you have enough teams that coordinating a single deploy becomes the bottleneck, splitting the system so teams can ship independently is a genuine win. Conway's law says your architecture will mirror your communication structure whether you like it or not, so you may as well design the boundaries deliberately.
But that logic runs in both directions. If you have three teams, you want roughly three deployable units, not thirty. Splitting a nine-person company into a dozen services does not give you team autonomy; it gives every engineer part-time ownership of four services none of them understand deeply. The boundary that helps a 200-person org actively hurts a 15-person one. The right number of services is downstream of your org chart, and your org chart is small.
Microservices are a solution to a people problem that pretends to be a solution to a technology problem. If you do not have the people problem yet, you are buying the cure for a disease you do not have, and the cure has serious side effects.
The modular monolith I keep recommending
What I push teams toward now is the boring middle: a single deployable application with hard internal boundaries. Modules with their own schemas, their own public interfaces, and a rule that nothing reaches across a boundary except through the published contract. In .NET this is unglamorous and effective. You get clean seams without giving up transactions, in-process calls, or the ability to run the whole thing on a laptop.
The quiet benefit is that a modular monolith keeps your options open. If the payments module genuinely needs to scale independently in two years, the boundary is already there and you extract it then, with real load data telling you where the seam should be. You get the design discipline of microservices with none of the operational tax, and you defer the irreversible decision until you actually know something. That is the opposite of how most premature splits happen, which is on a whiteboard, from a guess.
The latency nobody budgets for
Here is a number that changed how I think. An in-process method call in a .NET app is measured in nanoseconds. A gRPC call to a service in the same availability zone is, realistically, a couple of milliseconds once you count serialization, the network, and deserialization. HTTP with JSON is worse. That is a difference of roughly a million-fold per call.
For one call, who cares. But that authorization I mentioned touched eleven services, several of them more than once, and the p99 latency of the whole flow was north of 600 milliseconds for what should have been a sub-50ms operation. Customers felt it at checkout. We had not engineered a slow system on purpose. We had assembled it out of individually reasonable calls, and the latency was an emergent property of the shape. You cannot profile your way out of a problem that lives in the architecture.
When the split actually is right
I want to be fair, because there are real cases and I have built them. Extract a service when you have a component with a genuinely different scaling profile, a fraud-scoring engine that needs GPUs and bursts to ten times baseline while the rest of your app sits flat. Extract when a compliance boundary demands physical isolation of certain data. Extract when a subsystem is written in a different language for good reason, or when a specific team's release cadence is truly incompatible with everyone else's and you have measured that it is the bottleneck.
Notice the shape of every one of those. They are specific, load-bearing reasons, discovered from operating the system, not predicted from an architecture diagram. The test I apply is simple: can you name the concrete pain this split removes, in a sentence, without using the words "scale" or "clean"? If you cannot, you are not solving a problem. You are decorating one.
If you cannot run a monolith, you cannot run this
The uncomfortable truth I have watched play out repeatedly is that microservices demand a level of operational maturity that most teams reaching for them do not have. If your monolith has flaky tests, no real observability, and a deploy process that involves someone holding their breath, distributing it does not fix any of that. It multiplies it. Now you have twelve deploy processes and twelve chances to hold your breath.
Distributed systems are strictly harder to operate than centralized ones. They ask for mature CI/CD, contract testing, distributed tracing, service discovery, and an on-call culture that can reason about partial failure at 3am. A team that has not yet earned its first monolith is not ready to run a fleet. The order of operations matters, and it is not the order the conference talks imply.

Conclusion
The most expensive architecture decisions are the ones that feel most like good engineering while you are making them. Premature microservices feel responsible. They feel like planning ahead. That is exactly what makes them dangerous, because nobody in the review will argue against separation of concerns, and so the decision passes on vibes and costs you for years. If I could give a new CTO one rule, it would be this: earn the right to split. Run the monolith until it hurts in a specific, measurable place, then cut exactly there and nowhere else. The teams I admire are not the ones with the most services. They are the ones who could tell you, to the service, why each one exists.