Back to blog

Connection Pool Exhaustion: The Real Fix

Database connection pool exhaustion causes hanging requests and timeouts under load. Here is what actually causes it and how to fix your pool setup for good.

0 views 5 min read
Share
Connection Pool Exhaustion: The Real Fix

Pool timeout exceeded while waiting for a connection. Or worse — no error at all, just requests that quietly hang for twenty or thirty seconds before timing out somewhere downstream. That's connection pool exhaustion, and it's one of those production incidents that looks exactly like a database outage right up until you check the database and find it sitting at 4% CPU, doing nothing.

What Pool Exhaustion Actually Looks Like

Your app logs show almost nothing useful, because nothing has technically failed yet — the request is just waiting in line for a free connection. The database's own logs are where the real signal is: Postgres will log something like FATAL: remaining connection slots are reserved for non-replication superuser connections, or your pool library (pg, Sequelize, Knex) throws a TimeoutError after connectionTimeoutMillis runs out.

P95 latency is usually the first dashboard that moves. Average latency can look fine for a while because most requests still get a connection quickly — it's the unlucky tail that queues behind a handful of connections that never got released.

Why Pools Run Out

A connection pool caps how many open database connections your app is allowed to hold at once. Every request that touches the database checks one out, uses it, and checks it back in. Exhaustion happens one of two ways.

The boring way: real traffic exceeds what the pool was sized for, and requests queue because there's genuinely more concurrent demand than connections available. This is rare in practice — most apps never get close to their configured max.

The common way: connections get checked out and never checked back in. A client is grabbed, an error gets thrown partway through the query, and the code that was supposed to release it never runs. The pool slowly bleeds connections until none are left — this is the actual connection pool exhaustion pattern behind most "random" production slowdowns, and it looks identical to a traffic spike even on a quiet Tuesday afternoon.

The Mistake Most Teams Make First

Most advice here tells you to just raise max in your pool config. I'd skip that as the first move. Bumping the ceiling doesn't fix a leak, it just buys you a few more minutes before the same bug exhausts the bigger pool — and it can make things worse. If your Postgres instance has max_connections set to 100 and you're running five app instances each with a pool of 50, you'll blow through the database's own connection ceiling before any single app instance even notices it's struggling.

Sizing the pool matters, but it's the second fix, not the first. The first fix is finding where a connection is being checked out and not released.

Finding the Actual Leak

Before touching config, check whether connections are actually leaking or whether you're genuinely under-provisioned:

  • Query pg_stat_activity and count connections by application_name or state — a growing number of idle in transaction connections over time is the leak signature.
  • Add a log line (even a temporary one) at every pool.connect() call and its matching .release() — if the counts drift apart under load, something isn't releasing.
  • Check every code path that calls client.query() directly without a surrounding try/finally — that's almost always where the leak lives.

I debugged this exact scenario on a client's Node API last year. One endpoint called pool.connect() and ran a query with no try/finally around it at all. Under normal traffic, validation errors were rare enough that nobody noticed. Then a marketing email went out, traffic tripled, validation errors got more common because of malformed form submissions, and the API fell over in about four minutes flat.

The Fix: Sizing and Releasing Correctly

Fixing connection pool exhaustion for good means two things need to be true at the same time: the pool has to be sized against the database's real ceiling, not a guess, and every single connection checkout has to be guaranteed a release, even when the query throws.

js
156 tokens
const { Pool } = require('pg');
const pool = new Pool({
max: 15, // size against Postgres's max_connections,
// divided across all app instances — not traffic volume
idleTimeoutMillis: 10000, // release idle clients back to the pool fast
connectionTimeoutMillis: 3000, // fail fast instead of queuing forever
});
async function getOrderById(id) {
const client = await pool.connect();
try {
const { rows } = await client.query(
'SELECT * FROM orders WHERE id = $1',
[id]
);
return rows[0];
} finally {
client.release(); // the line almost every leak is missing
}
}

client.release() inside finally is the actual fix — it runs whether the query succeeds, throws a database error, or throws a validation error before the query even executes. If your codebase has more than one or two places calling pool.connect() directly, that pattern is worth wrapping in a small helper so nobody can forget the finally block by accident.

Frequently Asked Questions

Is it safe to just set max to something huge like 200?

No, not on its own. A higher max only helps if your database can actually support that many concurrent connections — check max_connections on the Postgres side first, and remember that number has to be shared across every app instance, not just one.

Why does Postgres show "too many connections" but my app never throws a timeout error?

Because something else connected directly to the database outside your pool — a migration tool, an admin script, a second service, or a monitoring agent — and used up slots your pool never knew about. Your pool's own max only limits connections it manages itself.

Key Takeaways

Pool exhaustion is almost never a sizing problem on the first pass — it's a missing finally. Fix the leak before you touch max, size the pool against the database's real connection ceiling divided across every app instance that shares it, and wrap every manual pool.connect() call in a helper that can't forget to release.

Found this useful? Share it.

Share