A few years ago one of my teams shipped a feature that felt trivial in review: when a payment settled, we wrote the settlement record to Postgres and then published an event to Kafka so the reconciliation service could pick it up. Two lines of code, one after the other. It passed every test. It ran fine for eight months.
Then a broker rebalance caused a two-second publish timeout right after the database commit succeeded. The row was in Postgres. The event never reached Kafka. Reconciliation had no idea the payment existed, and for eleven minutes our ledger and our downstream systems disagreed about roughly forty thousand euros. Nobody lost money, but I spent a very unpleasant afternoon explaining to a very calm auditor why our two sources of truth were not, in fact, in agreement. That is the dual-write problem, and I have watched it bite good engineers over and over.
What the dual-write problem actually is
A dual write is any operation where you change two systems that cannot participate in the same transaction, and you need both changes to happen or neither. Write to the database and publish to a message bus. Write to the database and call a third-party API. Update your own store and update a cache. Update two microservices' databases in one request handler. The shapes vary, but the disease is the same: you have two commits, and there is a window between them where a crash, a timeout, or a deploy leaves you in a state that should be impossible.
The reason this is so easy to get wrong is that the happy path is boring and correct. The code reads like a to-do list. Save the order, then notify the warehouse. Ninety-nine point nine percent of the time both succeed in the order you wrote them, and you never think about it again. The failure lives entirely in the gap, and the gap only opens under load, during incidents, or on the one Tuesday your network hiccups. So the bug ships, sleeps, and wakes up at the worst possible moment.
The tempting fixes that don't work
When people first meet this problem, they reach for ordering. "Just publish the event first, then commit the database." Now you can emit an event for a payment that never actually persisted, which is arguably worse because you have told the rest of the world something false. Flip it back and you are where I started: committed data, lost event.
The next instinct is retries. Wrap the publish in a retry loop, maybe with exponential backoff, and call it resilient. Retries genuinely help with transient blips, but they do not close the gap. If the process is killed between the commit and the successful publish, no in-memory retry loop survives to finish the job. A pod gets OOM-killed, a deploy rolls the container, the host reboots. The retry state was in memory, and memory is gone. You have made the window smaller, not closed it, and a smaller window is more dangerous because it lulls you into thinking the problem is solved.
The dangerous version of this bug is not the one that fails loudly in QA. It is the one that works perfectly until the exact moment your infrastructure is already on fire, and then quietly corrupts state while everyone is looking at the fire.
Two-phase commit, and why I mostly avoid it
The textbook answer is a distributed transaction: two-phase commit, XA, a transaction coordinator that makes both systems agree to commit or both roll back. On paper it is exactly what you want. In practice I have never once been glad I reached for it in a payments system. Kafka does not do XA in any way you want to rely on. Most modern HTTP APIs certainly do not. And even when both sides technically support it, 2PC introduces a coordinator that can itself fail during the prepared-but-not-committed phase, leaving locks held and resources blocked until someone intervenes.
There is a deeper reason I stay away. Two-phase commit couples the availability of your systems together. If your message bus is slow, your database transactions now hold locks longer, and a problem in one component becomes a problem in all of them. In a regulated environment where I care enormously about the ledger staying available and consistent, I do not want its health tied to a broker's mood. So I treat 2PC as a last resort for the rare case where I genuinely control both resource managers and can tolerate the coupling. That is almost never.
The pattern I actually reach for: transactional outbox
The fix that has earned my trust is the transactional outbox. The insight is a little obvious once you see it: the only place you can atomically write two things is inside a single database transaction. So instead of writing to the database and then publishing, you write to the database and record the intent to publish in the same transaction, in an outbox table. One commit. Either both the business row and the outbox row land, or neither does. There is no gap.
A separate process then reads unpublished rows from the outbox and pushes them to the message bus, marking them done once the broker acknowledges. If that publisher crashes mid-flight, it simply re-reads the unpublished rows on restart and tries again. The event might be delivered more than once, but it will never be lost. You have converted an impossible-to-solve "exactly once across two systems" problem into a very solvable "at least once, plus idempotent consumers" problem.
BEGIN;
INSERT INTO settlements (id, payment_id, amount_cents, status)
VALUES (@id, @paymentId, @amount, 'SETTLED');
INSERT INTO outbox (id, aggregate_id, event_type, payload, created_at)
VALUES (
@eventId,
@paymentId,
'settlement.completed',
@jsonPayload,
now()
);
COMMIT;
-- A separate relay reads unpublished outbox rows and pushes to Kafka,
-- marking published_at only after the broker acknowledges.
How to drain the outbox
You have two honest ways to move rows from the outbox to the broker, and the choice matters more than people expect.
- Polling relay: a worker queries for rows where published_at is null, ships them in order per aggregate, and marks them done. Dead simple, easy to reason about, easy to run in any language. The cost is polling latency and load on the database, though a tight loop with a small batch size gets you sub-second latency without much strain.
- Change data capture: point Debezium or your database's logical replication at the outbox table so inserts stream straight to Kafka. Lower latency, no polling load, but you have added a meaningful piece of infrastructure with its own failure modes, its own upgrades, and its own on-call learning curve.
My default is a polling relay until throughput or latency genuinely forces the move to CDC. I have watched teams adopt Debezium on day one for a service doing thirty events a minute, then spend a quarter learning its operational quirks. Start boring. Earn your complexity. The polling version fits in a screen of code and I can explain it to a new hire in five minutes, which is worth more than a few hundred milliseconds most of the time.
Idempotency is not optional
The outbox guarantees at-least-once delivery, which means duplicates are not an edge case, they are a promise. If your consumer credits a merchant account every time it sees a settlement event, and the relay redelivers after a crash, you have just paid someone twice. I have cleaned up exactly that mess, and clawing money back from a merchant is a conversation nobody enjoys.
So every consumer needs a way to recognize an event it has already processed. The cleanest approach is a natural idempotency key carried in the event, checked against a processed-events table before you act. Insert the key, do the work, and let a unique constraint reject the duplicate.
INSERT INTO processed_events (event_id, consumer, processed_at)
VALUES (@eventId, 'ledger-service', now())
ON CONFLICT (event_id, consumer) DO NOTHING;
-- If zero rows were affected, we have seen this event. Skip the side effect.
Design your side effects so that replaying them is safe. "Set balance to X" is safer than "add Y to balance" whenever the domain allows it. When it does not, the processed-events guard is your seatbelt, and you wear it every time.
Ordering, poison messages, and the boring details
Two things will surprise you in production. First, ordering. If two events for the same account arrive out of order, your consumer can compute nonsense. Partition your topic by the aggregate id so all events for one account land on one partition and stay ordered. Do not partition by event type or, worse, round-robin, unless you enjoy debugging balances that flicker.
Second, poison messages. Sooner or later the relay hits a row it cannot publish or a consumer hits an event it cannot process, and a naive loop will retry that one row forever while everything behind it starves. You need a retry count, a dead-letter path, and an alert when something lands there. I have seen an entire outbox back up for six hours because one malformed payload jammed the head of the line and nobody was watching the queue depth. Instrument the age of the oldest unpublished row. That single metric is the best early warning you will get.
When it is fine to just not care
I am not going to pretend every dual write deserves an outbox. If you are writing a settlement to the ledger and firing an analytics event that feeds a dashboard nobody makes decisions from at three in the morning, losing that event costs you a slightly wrong chart. Building outbox machinery for it is over-engineering, and over-engineering is its own kind of technical debt.
The question I ask is blunt: if these two writes disagree, who gets hurt and how badly? If the answer touches money, regulatory reporting, or a customer's trust, the outbox is cheap insurance. If the answer is "a graph looks funny," ship the naive version and move on. Reserve your reliability budget for the writes that actually matter, because you do not have an infinite one and pretending otherwise just spreads your attention too thin to protect anything well.

Conclusion
The thing I wish someone had told me earlier is that the dual-write problem is not really a technology problem. It is a problem of pretending two systems are one. Every fix that works accepts that they are separate and builds a bridge you can actually inspect. The outbox does not make the gap disappear by magic; it moves the atomic decision into the one place you already have atomicity, and it makes the recovery path something you can read, test, and reason about at two in the morning. If you take one habit from all this, make it this: whenever you find yourself writing to two systems in a row, stop and ask where the crash goes. If you cannot answer, you have not finished the feature.