Settlement is the part of a payments business that nobody talks about and everybody fights about in production. It is the machinery that takes the noisy stream of authorizations, captures, refunds, chargebacks, and fees, and turns it into a single, defensible answer to a deceptively simple question: who gets paid what, and when. Get it right and it disappears. Get it wrong and you spend your quarters reconciling spreadsheets and apologizing to merchants.
I have designed and rebuilt settlement systems at two different companies, and the lessons rhymed both times. This is how I think about the problem now: the invariants I refuse to compromise on, the data model that holds up under audit, and the operational scaffolding that keeps the whole thing trustworthy. None of it is glamorous. All of it matters.
What Settlement Actually Is
People conflate authorization, clearing, and settlement constantly, so it is worth being precise. Authorization is a promise: the issuer agrees the funds exist and reserves them. Clearing is the exchange of transaction detail between parties. Settlement is the actual movement of money to discharge the obligations created by all the prior steps. In a card scheme this is the scheme paying the acquirer; in our world it is us paying the merchant their net proceeds after fees, reserves, and adjustments.
The critical insight is that settlement is not a single event. It is a continuous process of accruing obligations and then discharging them on a schedule. A transaction that authorized on Monday, captured on Tuesday, and partially refunded on Thursday produces a chain of obligations, and the settlement system has to compute the net position across that chain for a given merchant over a given window. That window is usually a calendar day in a specific timezone, which sounds trivial and is the source of an astonishing number of bugs.
Once you accept that settlement is the discharge of accrued obligations rather than the processing of individual transactions, the architecture follows. You are building a ledger first and a payout engine second.
The Ledger Is the Product
The single most important decision I make on a settlement system is to put a double-entry ledger at the center and make every other component read from or write to it. Not a balances table that gets mutated. A ledger of immutable, append-only entries where every movement of value has equal and opposite postings, and balances are derived by summing entries. This is centuries-old accounting practice for a reason: it makes errors visible instead of silent.
The temptation, especially early, is to keep a running balance column and update it in place because it is faster to query. Resist this. A mutable balance is a single point of corruption with no audit trail. When a merchant disputes their payout six weeks later, you need to replay exactly how the number was derived, and you cannot replay an UPDATE statement that overwrote the previous value.
If you cannot reconstruct any balance at any point in time by replaying immutable entries, you do not have a ledger. You have a cache that lies to you under load.
Performance objections are real but solvable. You materialize balances into snapshot rows on a schedule and treat them strictly as a derived cache that can be rebuilt from the entries at any time. The entries remain the source of truth. The snapshot is an optimization you can throw away and regenerate, which is exactly the property you want.
Modeling Money and Time
Two primitives cause more incidents than anything else: money and time. For money, never use floating point. Store minor units as integers and always carry the currency code alongside the amount. A bare integer of 1500 is meaningless; 1500 in JPY and 1500 in USD differ by two orders of magnitude in decimal places. I model money as a value type that refuses to perform arithmetic across currencies.
For time, the rule is to store everything in UTC and attach the business timezone as explicit context wherever a settlement window is computed. A merchant in Berlin settles on a Berlin calendar day, and the boundary between Tuesday and Wednesday is a property of their configuration, not of your servers. I have seen a payout reconciliation drift by a full day because someone assumed the database timezone matched the merchant's.
public readonly record struct Money
{
public long MinorUnits { get; }
public string Currency { get; }
public Money(long minorUnits, string currency)
{
if (string.IsNullOrWhiteSpace(currency) || currency.Length != 3)
throw new ArgumentException("Currency must be a 3-letter ISO code.");
MinorUnits = minorUnits;
Currency = currency.ToUpperInvariant();
}
public static Money operator +(Money a, Money b)
{
if (a.Currency != b.Currency)
throw new InvalidOperationException(
$"Cannot add {a.Currency} and {b.Currency}.");
return new Money(checked(a.MinorUnits + b.MinorUnits), a.Currency);
}
}
The checked arithmetic matters. Overflow on a high-volume merchant during a busy season is the kind of bug that produces a silently wrong payout, and silent wrongness is the enemy. I would rather throw an exception and halt a batch than emit a number nobody can trust.
The Settlement Batch Lifecycle
A settlement run is a batch with a strict state machine, and modeling those states explicitly is what makes the system operable. A batch moves through a defined progression, and the transitions are the only places where money-affecting decisions happen. Everything in between is read-only computation against an immutable snapshot of the ledger.
- Pending — the window has closed but the batch has not yet been computed.
- Calculating — net positions are being summed; the batch is locked to a ledger cursor.
- Ready — amounts are computed and held for review or automatic release.
- Submitted — payout instructions have been handed to the banking rail.
- Confirmed — the rail has acknowledged the funds movement.
- Failed — the rail rejected or returned the payout, requiring remediation.
The reason to make these states first-class is that real payouts fail, and they fail asynchronously. A bank transfer can be accepted, then returned three days later because the merchant closed their account. Your batch model has to accommodate a Confirmed batch transitioning to a return-handling flow without corrupting the ledger. We do this by never deleting or editing the original postings; the return is a new set of postings that reverses the effect, leaving the original history intact.
I also insist that the calculating step pin itself to a specific ledger position, a cursor, before it sums anything. If new entries arrive mid-calculation, they belong to the next batch, not this one. Without that pin, two runs over the same window can produce different totals depending on timing, and reproducibility is non-negotiable.
Idempotency and Exactly-Once Payouts
The scariest failure mode in settlement is paying a merchant twice. Distributed systems do not give you exactly-once delivery for free, so you have to engineer it into the boundary where instructions leave your system and enter the banking rail. The mechanism is an idempotency key derived deterministically from the batch identity, persisted before the call, and honored by both your code and the rail.
Concretely, I generate the idempotency key from the batch ID and the payout destination, store it in a table with a unique constraint, and only then submit. If the submission times out and the worker retries, the unique constraint refuses the duplicate insert, and we know to query the rail for the status of the existing instruction rather than sending a new one. The database constraint is the real guard; the application logic is just polite.
CREATE TABLE payout_instruction (
id BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
batch_id BIGINT NOT NULL REFERENCES settlement_batch(id),
merchant_id BIGINT NOT NULL,
idempotency_key TEXT NOT NULL,
amount_minor BIGINT NOT NULL,
currency CHAR(3) NOT NULL,
status TEXT NOT NULL DEFAULT 'submitted',
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT uq_payout_idempotency UNIQUE (idempotency_key)
);
This pushes correctness down to the layer least likely to lie: the database. Application code can be redeployed mid-flight, a queue can redeliver a message, a pod can be killed between the call and the commit. The unique constraint survives all of it. I trust constraints over code, because constraints cannot be skipped by a hotfix.
Fees, Reserves, and Adjustments
The gross amount a merchant transacts is rarely what they receive. Settlement is where fees, rolling reserves, chargeback holdbacks, and manual adjustments all collide, and every one of them must be a ledger entry rather than a deduction baked into a single number. If a merchant asks why their payout was lower than expected, the answer should be a list of itemized postings, not a hand-wave.
Reserves deserve special care because they are obligations that span time. A rolling reserve might withhold a percentage of each day's volume and release it ninety days later. That is two distinct movements: a debit into a reserve account now and a scheduled credit back to the merchant's available balance later. Modeling it as two postings against an internal reserve account means the money is always accounted for and never simply vanishes from the merchant's view.
Adjustments are the human escape hatch, and they are dangerous precisely because they are manual. I require every adjustment to carry a reason code, an actor, and a reference to whatever incident or ticket justified it, and I require dual control above a threshold. An adjustment is just a ledger entry like any other, which means it inherits the same immutability and auditability as automated postings, but the metadata around it is what keeps it from becoming a backdoor.
Reconciliation as a First-Class Citizen
A settlement system that cannot prove it is correct is worse than useless, because it gives you false confidence. Reconciliation is the daily practice of comparing what your ledger says against what the outside world says, and it has to be built in from the start, not bolted on after the first incident. There are two reconciliations that matter, and they are different.
The first is internal: every batch must balance to zero across its postings. Debits equal credits, always, with no exceptions and no rounding slop. If a batch does not balance, you halt and investigate before any money moves. The second is external: the sum you instructed the bank to pay must equal the sum the bank actually moved, matched against the statement or rail report they return. Discrepancies here are usually fee surprises or returns you have not yet ingested.
Reconciliation is not a report you run after something breaks. It is the continuous proof that nothing is broken, and the day you skip it is the day you stop knowing whether your numbers are real.
I treat unreconciled differences as incidents with owners and deadlines, not as noise to be tolerated. A persistent few-cent discrepancy is not harmless; it is a signal that some assumption in your fee math or rounding is wrong, and small wrongness compounds. Driving every break to zero is what separates a system you can stake a banking license on from one you merely hope is right.
Observability and Failure Handling
Because settlement runs on a schedule and touches money, the operational requirements are stricter than for most services. I want to know within minutes if a batch failed to compute, if a payout was rejected, or if today's volume deviates sharply from the trailing average, because an anomaly in volume is often the first sign of an upstream data problem feeding bad inputs into the run.
Every batch emits structured metrics: total gross, total fees, total net, count of payouts, and the time spent in each state. We alert on missing batches as aggressively as on failed ones, because a batch that silently never ran is more dangerous than one that loudly errored. Failure handling itself is deliberately conservative: when in doubt, the system stops and asks a human rather than guessing, because the cost of a wrong payout dwarfs the cost of a delayed one.
That conservatism is a values statement as much as a technical one. A payout that is a day late is an inconvenience you can explain. A payout that is wrong is a trust failure and potentially a compliance event. I design every ambiguous code path to fail toward halting and escalating, never toward proceeding on an assumption.

Conclusion
Designing a settlement system is, in the end, an exercise in earning trust through structure. The double-entry ledger gives you truth you can replay, immutable postings give you an audit trail that holds up under scrutiny, idempotency keys backed by database constraints keep you from paying twice, and relentless reconciliation proves the whole thing is real. None of these are clever. They are disciplined. The cleverness is in refusing to compromise on them when the deadline pressure arrives, because every shortcut here is a future incident with a merchant's money attached. Build the ledger first, make every movement a posting, prove it balances every day, and the settlement system stops being the thing you fear and becomes the quiet, dependable backbone it is supposed to be.
