Connection Pooling and Why It Matters

Most of the worst production incidents I have managed in fintech did not start with a dramatic failure. They started quietly, with a service that was perfectly...

Originally published onanselmfowel.com

Most of the worst production incidents I have managed in fintech did not start with a dramatic failure. They started quietly, with a service that was perfectly healthy at noon and unresponsive by half past noon, while CPU and memory graphs sat flat and innocent. Nine times out of ten, the culprit was the database connection pool, or rather our misunderstanding of it. Connection pooling is one of those topics that everyone has heard of and very few people have actually reasoned through to the level of detail that matters when money is moving and latency budgets are measured in single-digit milliseconds.

Connection Pooling and Why It Matters
Connection Pooling and Why It Matters

This post is my attempt to write down what I wish every engineer joining a payments or lending team already understood about pooling. It is not a tutorial on a specific library. It is a way of thinking about a shared, finite resource that sits between your application and your most important system of record, and about why getting it right is the difference between a platform that degrades gracefully and one that falls over the moment traffic gets interesting.

What a Connection Actually Costs

A database connection feels free because, in code, opening one is a single line. In reality it is one of the more expensive things your application does. Establishing a connection to PostgreSQL or SQL Server involves a TCP handshake, a TLS negotiation if you are doing things properly, authentication, and then server-side setup: the database forks or assigns a backend process or thread, allocates memory for it, and primes session state. On Postgres in particular, every connection is backed by a full operating system process, which is why a few thousand of them will quietly eat your server alive.

When I have profiled this on real systems, the cost of opening a fresh connection has consistently landed somewhere between a few milliseconds and several tens of milliseconds depending on TLS and network distance. That sounds small until you multiply it by the request rate of a payment authorization service. If every API call pays that tax, you have built a system whose dominant cost is not the query but the ceremony around the query. Pooling exists to amortize that ceremony: open a set of connections once, keep them warm, and hand them out to whoever needs them.

The mental model I encourage is to treat connections like seats in a small conference room rather than like air. There are only so many, they take real effort to set up, and if everyone grabs one and forgets to leave, the room fills and the next person waits in the hallway. Once you internalize that connections are scarce and costly, most pooling decisions become obvious.

How a Pool Actually Behaves

A connection pool is a managed collection of open connections plus a checkout discipline. When your code asks for a connection, the pool either hands you an idle one, opens a new one if it is below its maximum, or makes you wait until someone returns theirs. When you are done, you return the connection to the pool rather than closing it, and it goes back into the idle set ready for the next caller. The connection is reused, not recreated.

The subtlety that trips people up is the waiting. A pool has a maximum size for a reason, and once every connection is checked out, additional requests do not fail immediately. They queue, up to some timeout. This is usually the right behavior, but it means a slow database does not produce slow responses; it produces a growing queue of threads all blocked waiting for a connection that is not coming back fast enough. Latency does not rise linearly. It rises gently and then vertically, which is exactly the shape of those quiet incidents I described at the start.

A connection pool does not make a slow query fast. It makes a slow query everyone's problem, because the connection that query is holding is one fewer connection for the rest of your traffic.

Sizing the Pool Is Counterintuitive

The instinct of most engineers, when a service is under load, is to raise the connection pool maximum. More connections, more throughput, surely. This is almost always wrong, and it is wrong in a way that is genuinely counterintuitive until you sit with the math. A database server has a finite number of cores and a finite amount of disk and memory bandwidth. Beyond a certain point, adding concurrent connections does not increase work completed; it increases context switching, lock contention, and cache thrashing. Throughput goes down while everything feels busier.

The guidance I keep coming back to, originally popularized by the HikariCP project, is that a small pool serving a queue almost always beats a large pool. A formula that has held up well in practice for an OLTP workload is connections equal to the number of cores times two, plus the effective number of spindles. For a modern eight-core database instance on SSDs, that points at a pool of roughly seventeen to twenty connections, not the two hundred that nervous teams tend to configure.

When I review a sizing decision, I ask the team to consider these factors rather than picking a round number:

  • The number of CPU cores and the nature of the storage on the database server, since those set the real ceiling on useful concurrency.
  • The number of application instances multiplied by their pool sizes, because the database sees the sum, not your per-pod configuration.
  • The presence of any connection multiplexer such as PgBouncer in transaction mode, which changes the arithmetic entirely.
  • The mix of fast and slow queries, since a few long-running reports can starve a pool sized for quick transactions.

Leaks and the Connection That Never Comes Back

The single most common pooling bug I have seen in production is the leaked connection: a code path that checks out a connection and, due to an exception or a forgotten return, never gives it back. Each leak permanently shrinks the usable pool by one. The service runs fine for hours because the pool is large enough to absorb a slow drip, and then one busy afternoon the last connection leaks and every subsequent request times out at once. The graph looks like a cliff, and the root cause is a line of code that ran cleanly thousands of times before.

In .NET the defense is disciplined disposal, and the language gives us the tools to make it nearly automatic. The pattern below is unremarkable and that is exactly the point; the unremarkable code is the code that does not leak.

public async Task<decimal> GetAvailableBalanceAsync(Guid accountId, CancellationToken ct)
{
    // Opening here actually rents from the ADO.NET pool, it does not
    // establish a new physical connection if an idle one exists.
    await using var connection = new SqlConnection(_connectionString);
    await connection.OpenAsync(ct);

    await using var command = new SqlCommand(
        "SELECT available_balance FROM accounts WHERE id = @id", connection);
    command.Parameters.Add(new SqlParameter("@id", accountId));

    var result = await command.ExecuteScalarAsync(ct);

    // Dispose on the connection returns it to the pool; it does not
    // tear down the underlying socket. The using block guarantees that
    // return even if ExecuteScalarAsync throws.
    return result is decimal balance ? balance : 0m;
}

The comments in that snippet carry the lesson. In ADO.NET, opening and closing a connection are pool rent and return operations, not socket operations, provided pooling is enabled and the connection string is identical across calls. The moment someone bypasses the using block, or builds connection strings dynamically so that each one spawns its own pool, the guarantees evaporate quietly.

Why Fintech Raises the Stakes

In a content website, a connection pool problem produces a slow page and an annoyed reader. In a payments platform, the same problem produces a failed authorization at the point of sale, a stuck disbursement, or a reconciliation job that does not complete before the banking cutoff. The blast radius is regulatory and financial, not merely cosmetic. That changes how I think about pool configuration from a tuning exercise into a reliability control that deserves the same scrutiny as any other part of the money path.

There is also a correctness dimension that pure-software teams sometimes miss. Connections carry session state: the transaction isolation level, temporary tables, session variables, and the open transaction itself. If a connection is returned to the pool mid-transaction, or with altered session settings, the next caller can inherit state they never asked for. In a financial system, a leaked open transaction holding locks can stall settlement for every other account, and a misattributed session setting can produce a query result that is subtly wrong rather than loudly broken. Loudly broken is recoverable. Subtly wrong is what gets discovered in an audit.

For these reasons I treat the pool as part of the trust boundary. We monitor it, we alarm on it, and we test its failure modes deliberately, because a financial platform that cannot reason about its own connection behavior cannot honestly claim to reason about its own correctness.

Pooling in a Serverless and Microservice World

The classic pooling advice was written for a world of a few long-lived application servers. Modern architectures break several of its assumptions. With dozens of microservices, each running many replicas, each holding its own pool, the database sees the product of all those numbers. A modest pool of twenty per pod becomes two thousand connections to the database when you scale to a hundred pods, and the database notices long before you do.

Serverless functions are worse, because each invocation may spin up an isolated execution environment with no shared pool at all. A traffic spike that launches a thousand concurrent function instances can attempt a thousand simultaneous fresh connections, which is precisely the storm pooling was invented to prevent. This is why an external pooler such as PgBouncer or RDS Proxy stops being optional in these architectures. It sits between your fleet and the database and multiplexes a large number of client connections onto a small number of real server connections.

The trade-off is that transaction-mode multiplexing forbids certain features that rely on session state spanning multiple statements, such as prepared statements that persist across transactions or session-level advisory locks. I have watched a team adopt a transaction-mode pooler and then spend a week debugging why their prepared statements vanished. The pooler was doing exactly what it promised; the application had assumed a session model that no longer existed. Understanding what your pooling layer guarantees is not optional knowledge.

Timeouts, Retries, and Graceful Failure

A pool without sensible timeouts is a trap. The connection-acquisition timeout determines how long a request waits in the queue before giving up, and it should be tuned in concert with your overall request budget. If your API promises a response within two hundred milliseconds, a connection-wait timeout of thirty seconds is dishonest; it lets requests pile up far past the point where the client has already given up, consuming resources to produce answers nobody is waiting for anymore.

I prefer to fail fast under saturation. A request that cannot get a connection within a tight budget should return a clear, retryable error rather than block. Combined with a circuit breaker, this turns a cascading collapse into a contained, observable degradation. The system sheds load instead of drowning in it, and the dashboards show a clean spike in rejections rather than an ambiguous wall of timeouts that could mean anything.

Retries deserve the same discipline. Blind retries against an already-saturated pool are how a minor blip becomes an outage, because every retry consumes a connection slot that a first-attempt request could have used. Retries belong behind jittered backoff and a budget, and they should never retry into a circuit that is already open. The goal is always to make the failure mode boring and predictable, because boring and predictable is what you want at three in the morning.

Observability Is the Whole Game

You cannot manage what you cannot see, and pools are notoriously invisible until they are on fire. The metrics I insist on for every service that talks to a database are the number of connections in use, the number idle, the number of requests currently waiting to check out a connection, and the time spent waiting. That last one, the wait time, is the leading indicator. It rises before errors appear, which gives you a window to act before customers feel anything.

I have learned to alarm on pool exhaustion as a first-class signal rather than inferring it from downstream symptoms. When the wait queue grows and acquisition time climbs, that is the truth of the situation, and the slow API responses and timeouts are merely its shadows. A team that watches the shadow chases ghosts; a team that watches the pool sees the cause. Pair these metrics with slow-query logging on the database side and you can usually name the offending query within minutes rather than hours.

The most valuable habit I can recommend is to load-test the failure mode on purpose, in a controlled environment, before production does it for you. Drive traffic past the pool's capacity and watch how the system behaves. Does it fail fast and shed load, or does it accumulate a quiet queue that eventually topples? You want to discover the answer to that question on a Tuesday afternoon with a test harness, not on a Friday night with real customers and a regulator's reporting window closing.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

Connection pooling matters because it sits at the precise point where your application's appetite meets your database's finite capacity, and in a financial system that meeting point is part of the money path. The right pool is usually smaller than your instinct suggests, sized to the database rather than to the application, protected by tight timeouts and honest failure modes, and watched with metrics that reveal stress before it becomes failure. Get those things right and the pool disappears from your incident reports entirely, which is the highest praise infrastructure can earn. Get them wrong and you will keep meeting the same quiet noon-time outage, wearing a different costume each time, until you finally sit down and reason about the seats in the room.

Chat with us