Every distributed system fails. That is not a pessimistic statement; it is an operational one. In the payments platforms I have run, the question was never whether a downstream dependency would degrade or disappear, but how the rest of the system would behave in the seconds after it did. The difference between a controlled incident and a cascading outage almost always comes down to two patterns: circuit breakers and graceful degradation.
I want to talk about these patterns the way I think about them on a Tuesday afternoon when a settlement provider starts returning timeouts. Not as architecture diagrams, but as decisions you encode into code, configuration, and on-call runbooks. The goal is simple to state and hard to achieve: when something breaks, the blast radius should be small, the customer experience should degrade gracefully, and the engineers on call should be able to reason about what is happening.
Why Naive Retries Make Outages Worse
The instinct of most engineers, when a call fails, is to retry it. That instinct is correct for transient errors and catastrophic for systemic ones. The dangerous moment is when a dependency is not down but slow. Latency creeps up, your timeouts start firing, and your retry logic dutifully sends the request again. Now you have doubled or tripled the load on a service that was already struggling, across every instance of your fleet at once.
This is the retry storm. A fraud-scoring service we depended on once slowed from twenty milliseconds to two seconds during a database failover. Within ninety seconds, retries from our authorization path had buried it under five times its normal volume, and a recoverable blip became a fifteen-minute outage. The fix was not more capacity. It was teaching our callers to stop calling.
That is the entire premise of a circuit breaker. When a dependency is failing, the most useful thing you can do is fail fast and stop sending it traffic, giving it room to recover. You trade a small amount of correctness for a large amount of stability, and that trade has to be made deliberately rather than by accident.
The Anatomy of a Circuit Breaker
A circuit breaker is a small state machine wrapped around a call to a dependency. It has three states. In the closed state, calls pass through normally and the breaker counts failures. When failures cross a threshold it trips to open, and subsequent calls fail immediately without touching the dependency. After a cooldown it moves to half-open, allowing a limited number of trial calls through to test whether the dependency has recovered. If those succeed it closes again; if they fail it returns to open.
The subtlety is in the thresholds. A breaker that trips too eagerly fails calls that would have succeeded; one that trips too late provides no protection. I prefer a failure-rate threshold over a raw count, evaluated over a rolling window with a minimum sample size, so that a breaker does not trip on the first two failures during a quiet period at three in the morning.
Here is the shape I use in .NET, built on Polly, which I have run in production for years:
var breaker = new ResiliencePipelineBuilder<HttpResponseMessage>()
.AddCircuitBreaker(new CircuitBreakerStrategyOptions<HttpResponseMessage>
{
FailureRatio = 0.5, // trip at 50% failures
MinimumThroughput = 20, // need 20 calls in window
SamplingDuration = TimeSpan.FromSeconds(30), // rolling window
BreakDuration = TimeSpan.FromSeconds(15), // cooldown before half-open
ShouldHandle = new PredicateBuilder<HttpResponseMessage>()
.Handle<HttpRequestException>()
.Handle<TimeoutRejectedException>()
.HandleResult(r => (int)r.StatusCode >= 500),
OnOpened = args =>
{
_logger.LogWarning("Breaker OPEN for {Dep} after {Ratio:P0}",
"fraud-scoring", args.Outcome.Result?.StatusCode);
_metrics.Increment("breaker.opened", tags: "dep:fraud-scoring");
return default;
}
})
.Build();
Notice what the breaker handles and, just as importantly, what it does not. A 400 is the dependency telling you that your request is wrong; tripping on it is pointless. A 503 or a timeout is a signal about the dependency's health. Encoding that distinction precisely is most of the work.
Timeouts, Bulkheads, and the Full Resilience Stack
A circuit breaker on its own is necessary but not sufficient. It needs to sit inside a stack of complementary patterns, and the order matters. The patterns I insist on for any critical dependency:
- Timeouts that are aggressive and explicit. A call with no timeout is a resource leak waiting to happen; I would rather fail at 800 milliseconds than hold a thread for thirty seconds.
- Bulkheads that cap concurrency per dependency, so a slow service cannot consume the whole connection pool and starve healthy paths.
- Circuit breakers that trip on sustained failure and protect the downstream from retry storms.
- Retries with jittered exponential backoff, applied only to idempotent operations and only for genuinely transient errors.
- Fallbacks that define what the system returns when all of the above have failed.
The sequencing is deliberate. A retry should sit outside a breaker, so that the breaker counts the underlying failures and the retry respects an open circuit. A timeout sits inside, converting a slow call into a fast failure that the breaker can count. Get this order wrong and you end up with retries that hammer an open breaker, or breakers that never see the failures because the retry swallowed them. That is worse than having no patterns at all, because it creates false confidence.
Bulkheads deserve special mention in payments. If your authorization path and your reporting path share a connection pool, a degraded reporting database can take down your ability to authorize transactions. That is an unacceptable coupling. Isolating resources by criticality is one of the highest-leverage decisions you can make, and it costs almost nothing.
Graceful Degradation Is a Product Decision
Circuit breakers stop the bleeding. Graceful degradation decides what happens next. When a breaker is open and a dependency is unavailable, you have to answer a question that is fundamentally about product and risk, not engineering: what is the least-bad behavior?
This is where many teams stop thinking. The lazy answer is to return an error to the user. Sometimes that is correct, but often there is a better degraded mode, and choosing it requires a conversation with product, risk, and compliance rather than a unilateral engineering call.
When the fraud-scoring service is unavailable, do we decline all transactions, approve all of them, or fall back to a conservative rules-based model that approves low-risk transactions and holds the rest for review? Each answer has a different fraud loss, a different customer-experience cost, and a different regulatory posture. There is no universally correct choice, only a choice the business has made deliberately and documented.
The degraded behavior must be a decision made in advance, written down, and agreed across functions. In a regulated environment you cannot improvise risk posture during an incident; the engineer on call at two in the morning should be executing a policy, not inventing one. When the fallback approves transactions, finance needs the exposure cap. When it declines them, support needs the messaging. Graceful degradation is a cross-functional artifact that happens to be implemented in code.
Stale Data as a Deliberate Fallback
One of the most effective degradation strategies is serving stale data when fresh data is unavailable. If your currency-conversion service is down, serving a rate fifteen minutes old is often far better than failing, provided you understand the financial exposure and cap how old you will go. This is the difference between a cache as a performance optimization and a cache as a resilience mechanism.
The discipline is to make staleness explicit and bounded. I want to know the age of every cached value I might serve in a degraded mode, a hard limit beyond which stale is worse than absent, and a record that the system served stale data so reconciliation and audit can account for it. A silent fallback to stale data is a future incident; a logged, bounded, deliberate one is resilience.
For reference data that changes slowly, such as merchant configuration or routing tables, this is almost free. For pricing or risk data the tolerance is much tighter and the conversation with risk is mandatory. The architecture is the same; the parameters are a business decision.
Observability and the Half-Open Trap
A circuit breaker that you cannot observe is a liability. The most important thing to instrument is the state transition. Every time a breaker opens, closes, or moves to half-open, that event should be a metric and a log line with the dependency name attached. When an incident starts, the first question is always which breakers are open, and you should answer it from a dashboard in seconds rather than by grepping logs.
The half-open state is where subtle bugs hide. When a breaker goes half-open, it permits a small number of trial requests. If your fleet has fifty instances and each goes half-open at the same moment, you send a synchronized burst of trial traffic that re-trips the breaker and convinces you the dependency is still down when it has actually recovered. This is the thundering-herd problem applied to recovery. Jittering break durations across instances mitigates it.
I also insist on a manual override. There are times when you, the operator, know more than the breaker does: you have confirmed the dependency is healthy and want to force the circuit closed immediately rather than wait for the automatic cycle. Equally, you sometimes want to force a breaker open to shed load from a dependency you know is fragile. The breaker is a tool for the on-call engineer, not a replacement for them.
Testing Failure Before It Tests You
The uncomfortable truth about resilience patterns is that they only run during the exact moments you least want surprises. A circuit breaker that has never been exercised is a hypothesis, not a control. The fallback path runs rarely by definition, which makes it the path most likely to contain the untested bug, the missing null check, the configuration that was never wired up.
This is why fault injection has to be part of the engineering practice rather than an annual fire drill. I want integration tests that force a dependency to time out and assert that the breaker opens and the fallback returns the agreed response, a staging environment where we can flip a real breaker open, and in mature organizations, controlled failure injection in production. The cultural shift is to treat the degraded path as a first-class feature with its own acceptance criteria. When we started writing explicit tests for every fallback, we found that roughly a third of them did something other than what the runbook claimed. That gap, discovered in a test rather than an incident, is the entire return on the investment.
The Organizational Side of Resilience
None of this is purely technical. The hardest part of running resilient payments systems is keeping the decisions about degraded behavior current as the business evolves. A fallback policy agreed two years ago may no longer match the company's risk appetite, the regulatory landscape, or the product. Resilience configuration rots silently when no one owns it.
I assign explicit ownership of each critical dependency's resilience posture to a team, and require that degraded-mode decisions be reviewed alongside the threat model, not buried in code. The breaker thresholds, timeout budgets, staleness limits, and fallback policies live in one place that product, risk, and engineering all read. When an incident occurs, the review asks not only what failed but whether the degraded behavior matched policy, and whether that policy is still right.
The teams that do this well treat an open circuit breaker during an incident as a sign the system worked. The breaker did its job. The customer saw a degraded but coherent experience. The on-call engineer executed a known policy. That is what success looks like, and it is very different from the heroics of debugging a cascading outage at four in the morning.

Conclusion
Circuit breakers and graceful degradation are often introduced as defensive libraries you bolt onto a service, and that framing undersells them. They are the mechanism by which a system makes deliberate, documented choices about how it behaves under stress, choices that span engineering, product, risk, and compliance. The breaker handles the mechanics of failing fast and protecting downstreams; graceful degradation answers the human question of what the system should do instead. Get the first right and you stop cascading outages. Get the second right and your customers barely notice. Observe and test both, give them clear owners, revisit their decisions as the business changes, and you turn the inevitability of failure into something you can actually live with.
