Notifications look like a small feature until you operate one at scale. Then you discover that the humble "we just send the customer an email" requirement is actually a distributed systems problem wearing a trench coat. In a regulated fintech business, a notification is not a nicety. It is the artifact a customer points to when they dispute a charge, the record an auditor asks for when reconstructing a payment lifecycle, and the message that has to arrive exactly once even when three services upstream retried the same event.
Over the last few years my teams have rebuilt our notification platform twice. The first version was a thin wrapper around an email provider, bolted directly onto the payments service. It worked until it did not. What follows is the architecture we settled on, the trade-offs that shaped it, and the operational lessons that only show up after you have sent a few hundred million messages.
Why Notifications Are a Distributed Systems Problem
The naive mental model is request-response: something happens, you call SendGrid or Twilio, the message goes out. That model breaks the moment you care about reliability. Your provider has an outage. Your own database commits but the subsequent send call times out. A consumer crashes after sending but before recording that it sent. Each of these is a partial failure, and partial failures are the entire discipline of distributed systems.
Once you accept that notifications are events flowing through an unreliable network of services, the right primitives become obvious. You need durable queues, idempotency keys, retries with backoff, dead-letter handling, and an audit trail. None of these are exotic. The hard part is composing them so that the common case is fast and the failure case is correct, rather than the other way around.
The second thing to internalize is that notifications are inherently multi-channel and multi-tenant. The same underlying event, say a settlement completing, might fan out into an email to the merchant, an SMS to a configured operations contact, a webhook to the merchant's own backend, and an entry in an in-app inbox. Treating each channel as a special case leads to a tangle. Treating them as variations on a single delivery contract keeps the system comprehensible.
Separating the Event from the Message
The single most useful decision we made was to draw a hard line between an event and a notification. An event is a fact about the world: payment.settled, kyc.review_required, statement.ready. It carries data and nothing about presentation. A notification is the decision to tell a specific recipient about that fact through a specific channel using a specific template. Upstream services emit events; they have no idea whether anyone will be notified.
This separation pays off constantly. Product wants to add a new SMS alert for failed direct debits? That is a routing and template change inside the notification service, with zero deployments to the payments domain. Compliance wants to suppress marketing-adjacent messages for a class of accounts? That is a policy applied at the notification layer. The services that own the truth stay clean, and the messy, frequently-changing presentation logic lives in one place that is built to absorb churn.
The teams that own business truth should never know how a customer gets told about it. The day your payments service imports an email templating library is the day you have lost the boundary.
The Ingestion and Queueing Layer
Events arrive on a message broker. We use a partitioned log for the high-volume streams and a classic queue for lower-volume control-plane events. The notification service subscribes, and the first thing it does on receiving an event is persist it. Not process it, persist it. The inbound event is written to a table with its broker offset and a derived idempotency key before any routing logic runs. This gives us a replayable record and a place to deduplicate.
The reason persistence comes first is that brokers give you at-least-once delivery, which means you will see the same event twice. If your idempotency check is the very first gate, duplicates are cheap to discard. If it comes later, you risk sending two emails for one settlement, and there is no apology that fully repairs a customer who got two conflicting balance alerts at 3am.
From there, routing expands one event into zero or more delivery intents, each targeting a channel and recipient. Those intents go onto per-channel queues. Separating queues by channel matters because channels fail independently and have wildly different throughput and latency characteristics. An email provider degradation should never back up your in-app inbox writes, and a slow third-party webhook endpoint should not starve SMS.
Idempotency and Exactly-Once Semantics
Exactly-once delivery is a phrase that gets thrown around loosely. In the strict sense it is impossible across an unreliable network. What you can build is exactly-once effect: the message may be attempted many times, but the recipient sees it at most once, and at least once in the happy path. The mechanism is an idempotency key carried end to end and enforced at the point of side effect.
We derive the key deterministically from the event identity and the delivery intent, so the same event reprocessed produces the same key. Before handing a message to a provider, we insert a row claiming that key inside the same transaction that records the send attempt. A unique constraint does the heavy lifting. If the insert fails, someone else already owns this send and we stop.
-- Claim an idempotency key before dispatching.
-- The unique index on idempotency_key makes a duplicate
-- claim fail loudly instead of producing a second send.
INSERT INTO notification_dispatch (idempotency_key, channel, recipient, status, created_at)
VALUES (@key, @channel, @recipient, 'claimed', now())
ON CONFLICT (idempotency_key) DO NOTHING
RETURNING id;
-- If RETURNING yields no row, another worker owns this dispatch.
-- We treat that as success and acknowledge the message,
-- because the effect (a single send) is already guaranteed.
The subtlety is the window between claiming the key and confirming the provider accepted the message. If we crash there, the row sits in claimed and we do not know whether the provider sent it. We resolve this with a reconciliation job that queries provider status APIs by our key, which we always pass as the provider-side idempotency token where supported. Belt and suspenders, but at this volume the belt and the suspenders both earn their keep.
Channel Adapters and the Provider Abstraction
Every channel sits behind an adapter that implements one interface. The notification core does not know it is talking to a specific email vendor; it knows it is talking to something that accepts a rendered message and returns an accepted, rejected, or transient-failure result. This abstraction is what lets us run two email providers concurrently and shift traffic between them when one degrades, without touching routing or templating.
Concretely, the contract looks like this in our .NET codebase. Keeping it deliberately small forces vendor-specific quirks to stay inside the adapter rather than leaking into the core.
public interface IChannelAdapter
{
string Channel { get; }
// Returns a result the core can act on without knowing the vendor.
// Transient failures are retried; permanent ones are dead-lettered.
Task<DispatchResult> SendAsync(RenderedMessage message, CancellationToken ct);
}
public sealed record DispatchResult(
DispatchOutcome Outcome, // Accepted, Rejected, Transient
string? ProviderMessageId,
string? Reason);
A clean adapter boundary also makes failure classification consistent. Each adapter is responsible for mapping the chaos of vendor responses (HTTP 429, malformed-recipient errors, soft bounces, hard bounces) into our three outcomes. The core then applies one retry policy regardless of channel. When a new vendor enters our stack, the only question we have to answer is how their errors map onto those three buckets.
Retries, Backoff, and the Dead-Letter Path
Transient failures get retried with exponential backoff and jitter. The jitter is not optional. Without it, a provider outage that recovers at a fixed moment produces a thundering herd as every backed-off message retries simultaneously, and you knock the recovering provider straight back down. Jitter spreads the load and is two lines of code; skipping it is a self-inflicted incident.
We cap retries, and we cap total time-in-flight separately. A message that has been bouncing around for an hour is rarely still worth sending. Many notifications are time-sensitive: a one-time passcode is useless after ninety seconds, a fraud alert loses most of its value within minutes. We attach a relevance deadline to time-critical intents so they expire rather than arriving uselessly late.
- Exponential backoff with full jitter, capped at a per-channel maximum interval.
- A retry ceiling and an independent wall-clock deadline, whichever trips first.
- A relevance window for time-sensitive messages, after which the intent is dropped and logged rather than delivered.
- A dead-letter queue for anything that exhausts retries, with the full event and rendering context attached.
- Alerting on dead-letter depth, because a rising floor there is the earliest signal that a provider or template is broken.
The dead-letter path is where operational maturity shows. Messages that fail permanently are not silently discarded; they land in a queue with enough context to be replayed once the underlying problem is fixed. After a provider incident we can selectively reprocess only the messages that were still within their relevance window, which keeps us from re-sending a flood of stale alerts.
Rendering, Templates, and Localization
Rendering is the step where an event plus a template plus recipient context becomes the literal bytes we hand to a channel. We treat templates as versioned, reviewable artifacts, not strings in a database that anyone can edit live. Every template change goes through the same review and deployment discipline as code, because a malformed template that omits a legally-required disclosure is a compliance incident, not a typo.
Localization and personalization live here too, and they introduce a quiet category of bug: a template that renders fine in English silently breaks when a German string is forty percent longer, or a currency formats with the wrong separators for a locale. We render every template against a fixture set for each supported locale in CI, and we snapshot the output. A diff in the snapshot is a deliberate decision, never an accident discovered by a customer.
One more rule we enforce: rendering is pure and deterministic. Given the same event and template version, it produces identical output. That determinism is what makes the idempotency story hold and what lets us reproduce exactly what a customer received when they call to dispute it. If rendering reached out to a live service for data, we would lose both properties, so any data a template needs is resolved upstream and frozen into the delivery intent.
Observability, Auditability, and Compliance
In a regulated environment, being able to answer "what did we send this customer, when, and why" is not a feature you add later. We carry a correlation identifier from the originating event through every queue, adapter, and provider call, and we record an immutable audit entry at each transition: event received, intent created, dispatch claimed, provider accepted, delivery confirmed or bounced. This is queryable and retained per our regulatory obligations.
The audit trail also doubles as our observability backbone. The same transitions that satisfy an auditor feed dashboards on delivery latency per channel, provider acceptance rates, bounce rates by template, and end-to-end time from event to confirmed delivery. When a merchant reports they never got a settlement notice, support can trace the exact path in seconds and tell whether we never sent it, the provider rejected it, or it bounced at the recipient's mail server.
Privacy is the constraint that shapes how much of this we keep. We log identifiers and metadata aggressively but treat message bodies as sensitive: they may contain balances, names, and partial account details. Bodies are retained only as long as compliance requires, encrypted at rest, and access to them is itself audited. The discipline is to log enough to reconstruct the truth without turning your log store into an unguarded copy of customer financial data.
Scaling the Platform Without Scaling the Pain
Scaling this architecture is mostly a matter of letting the seams we already built do their job. Because work is partitioned by channel and by tenant, we scale consumers horizontally and independently. A surge of statement-ready events at month-end spins up more email workers without touching the SMS path. Back-pressure is handled by the queues themselves; when a downstream provider slows, the queue grows, our autoscaling and alerting react, and nothing upstream falls over.
The harder scaling problems are not throughput, they are blast radius and noisy neighbors. One misconfigured high-volume tenant can saturate a shared provider account and degrade everyone. We enforce per-tenant rate limits and fairness in scheduling so that a single large merchant cannot starve the long tail. We also keep the ability to pause a single event type or tenant without halting the platform, which has turned several would-be incidents into a quiet config change.
If I had to compress the scaling philosophy into one line, it is this: design so that the failure of any one channel, provider, or tenant is contained and observable, then let commodity horizontal scaling handle the volume. The expensive engineering goes into the boundaries, not the throughput.

Conclusion
A notification service that scales is not built by finding a clever trick; it is built by taking ordinary distributed systems primitives seriously and applying them with discipline. Separate events from notifications. Persist before you process. Make idempotency a first-class, end-to-end concern. Hide vendors behind a narrow adapter, retry with jitter, and dead-letter what you cannot send. Treat rendering as deterministic, versioned, and reviewable. And in a regulated business, build the audit trail first, because the day a regulator or a customer asks what you sent, "we are not sure" is not an answer you can afford. Do those things and the system will quietly carry hundreds of millions of messages, which, after all, is the only kind of notification system worth running.
