
MongoDB Connection Pooling: Avoiding Too Many Connections
It usually starts with an alert. Atlas says you're at 90% of your connection limit, or your app logs fill with MongoServerSelectionError and pool timeout errors, or a deploy that added two more replicas somehow tipped the database over. The queries themselves are fast. There are just far too many connections, and each one costs memory and CPU on the server.
The fix is almost never "buy a bigger cluster". It's understanding how connection pooling works in MongoDB drivers, and then finding the code path that creates more clients (and more pools) than you think. Once you can do the math on how many connections your fleet should open, a connection storm becomes easy to diagnose.
This guide covers how driver pools work, the settings that control them, how to calculate expected connections, how to find where connections are coming from, and the patterns that cause connection explosions in real apps.
How a Driver Connection Pool Works
Every official MongoDB driver implements the same pooling specification, so the behavior is consistent across Node.js, Python, Java, Go, C#, and the rest.
When you create a MongoClient, the driver:
- Discovers the topology (every member of the replica set, or every
mongosin a sharded cluster). - Opens monitoring connections to each server. Modern drivers use a streaming heartbeat plus a separate round-trip-time connection, so expect about two monitoring connections per server per client.
- Creates a connection pool per server for application operations. Pools start empty (unless you set a minimum) and grow on demand.
When your code runs a query, the driver checks out a connection from the pool for the target server, sends the command, waits for the reply, and returns the connection to the pool. A connection is only "busy" for the duration of one operation, so a pool of 10 connections can easily serve thousands of requests per second when queries are fast.
If every connection in the pool is checked out, new operations wait in a queue until one is returned, or until a timeout fires. That wait is the pool doing its job: it's what stops your app from opening unlimited sockets.
The Settings That Matter
These options exist in every driver, usually as connection string parameters:
| Option | Default | What it controls |
|---|---|---|
maxPoolSize | 100 | Max connections per server, per client |
minPoolSize | 0 | Connections kept open even when idle |
maxIdleTimeMS | 0 (no limit) | Close connections idle longer than this |
maxConnecting | 2 | Max connections being established concurrently per pool |
waitQueueTimeoutMS | varies by driver | How long an operation waits for a free connection |
connectTimeoutMS | 30000 | TCP connect and handshake timeout |
They go in the connection string:
mongodb+srv://app:secret@cluster0.example.mongodb.net/shop?maxPoolSize=20&minPoolSize=2&maxIdleTimeMS=60000&appName=checkout-api
Or in client options, for example in Node.js:
import { MongoClient } from "mongodb";
const client = new MongoClient(process.env.MONGODB_URI, {
appName: "checkout-api",
maxPoolSize: 20,
minPoolSize: 2,
maxIdleTimeMS: 60_000,
waitQueueTimeoutMS: 5_000,
});
A few notes on these:
maxPoolSizeis per server. Against a three-member replica set, a client could in theory hold up tomaxPoolSizeconnections to each member. In practice, with the defaultprimaryread preference, almost all application connections go to the primary.maxConnectingsmooths out bursts. Without it, a traffic spike on a cold pool would try to open dozens of connections at once, each doing a TLS handshake and authentication, which is expensive for the server.maxIdleTimeMSis useful for bursty workloads. It lets the pool shrink after a spike instead of holding peak connections forever.appNameisn't a pool setting, but it's the single most useful option for debugging. It shows up in server logs and in$currentOp, so you can tell which service owns which connections.
Doing the Math
Here's the formula for the number of connections your fleet can open against the primary:
connections = processes x clients_per_process x (maxPoolSize + monitoring)
Where processes is every running process that creates a client: containers, pods, Gunicorn or PM2 workers, cron jobs, background workers.
A realistic example:
- 12 Kubernetes pods for the API
- each running Node.js in cluster mode with 4 workers
- one client per worker, default
maxPoolSizeof 100
That's 48 processes, each able to open up to 100 connections to the primary, for a ceiling of 4,800 application connections plus roughly 100 monitoring connections per server. You won't hit the ceiling every day, but a slow query that makes operations back up will push every pool toward its max at the same moment. That's how a small slowdown becomes a connection storm.
Now compare that with your cluster's limit. Atlas sets a maximum number of connections per tier; the free M0 tier allows 500, and dedicated tiers scale up with instance size (check the Atlas documentation for your tier's current figure). Self-hosted mongod has a configurable net.maxIncomingConnections, but the practical limit is memory: each connection uses around 1 MB of server RAM for its thread stack and buffers.
Sizing the Pool
For a single process, you can estimate the pool size you actually need with Little's law:
needed connections = requests per second x average operation time (seconds)
A worker handling 200 operations per second with a 5 ms average operation time needs about 200 x 0.005 = 1 connection on average. Give it headroom for bursts and slow queries and a pool of 10 to 20 is generous. The default of 100 is sized for a single large process, not for dozens of small ones.
A practical rule: the more processes you run, the smaller each pool should be. Many small workers with modest pools beat a few with the default of 100.
Finding Where Connections Come From
When connection counts look wrong, start on the server. In mongosh, get the totals:
db.serverStatus().connections;
{
current: 1842,
available: 49358,
totalCreated: 918233,
rejected: 0,
active: 37,
threaded: 1842,
exhaustIsMaster: 0,
exhaustHello: 61,
awaitingTopologyChanges: 0
}
Two things jump out here. current is 1,842 but only 37 are active, so the vast majority are idle, which points at oversized pools or too many clients. And totalCreated is huge, which means connections are being opened and closed constantly. That's the signature of code that creates a new client per request.
Next, group connections by the application that opened them:
db.getSiblingDB("admin").aggregate([
{ $currentOp: { allUsers: true, idleConnections: true } },
{
$group: {
_id: {
app: "$appName",
client: { $arrayElemAt: [{ $split: ["$client", ":"] }, 0] },
},
connections: { $sum: 1 },
},
},
{ $sort: { connections: -1 } },
{ $limit: 10 },
]);
[
{ _id: { app: "report-worker", client: "10.0.4.17" }, connections: 1204 },
{ _id: { app: "checkout-api", client: "10.0.3.22" }, connections: 41 },
{ _id: { app: "checkout-api", client: "10.0.3.23" }, connections: 39 },
];
Now you know which service and which host to look at. If you see connections with an empty or generic appName (like "MongoDB Shell" or a default driver string), start setting appName on every client. On Atlas, the Real-Time Performance Panel and the connections metric in the cluster monitoring view give you the same picture over time.
Watching the Pool from the Client Side
Drivers emit pool events you can log or turn into metrics. In Node.js:
const client = new MongoClient(uri, {
appName: "report-worker",
maxPoolSize: 10,
});
let open = 0;
client.on("connectionCreated", () => {
open += 1;
console.log(`pool: created (open=${open})`);
});
client.on("connectionClosed", (e) => {
open -= 1;
console.log(`pool: closed reason=${e.reason} (open=${open})`);
});
client.on("connectionCheckOutFailed", (e) => {
console.warn(`pool: checkout failed reason=${e.reason}`);
});
client.on("connectionPoolCleared", () => {
console.warn("pool: cleared (server marked unknown)");
});
If you see connectionCreated logged for nearly every request, the pool isn't being reused. If you see connectionCheckOutFailed with a timeout reason, the pool is too small for your concurrency or operations are too slow and backing up.
PyMongo, the Java driver, and the Go driver expose equivalent listeners (ConnectionPoolListener in PyMongo and Java, event.PoolMonitor in Go).
The Patterns That Cause Connection Explosions
Creating a Client per Request
By far the most common cause:
// Don't do this
app.get("/orders/:id", async (req, res) => {
const client = new MongoClient(process.env.MONGODB_URI);
await client.connect();
const order = await client
.db("shop")
.collection("orders")
.findOne({ _id: req.params.id });
res.json(order);
// no close(): leaked pool, monitoring connections, and timers
});
Every request builds a new pool and new monitoring connections, and without close() they linger. Even with close(), you pay a TCP connect, TLS handshake, and authentication on every request, which adds tens of milliseconds and burns server CPU. The fix is to create the client once at module level and reuse it:
// db.js
import { MongoClient } from "mongodb";
export const client = new MongoClient(process.env.MONGODB_URI, {
appName: "orders-api",
maxPoolSize: 20,
});
export const db = client.db("shop");
The Node.js driver connects automatically on the first operation, so you don't need an explicit connect() call, though calling it at startup is a good way to fail fast.
Dev Servers That Reload Modules
Frameworks with hot module reloading (Next.js dev mode, Vite, nodemon-style reloaders) can re-execute your db.js on every change, creating a new client each time while the old one stays alive. Cache the client on globalThis in development:
const globalForMongo = globalThis;
export const client =
globalForMongo._mongoClient ??
new MongoClient(process.env.MONGODB_URI, { maxPoolSize: 10 });
if (process.env.NODE_ENV !== "production") {
globalForMongo._mongoClient = client;
}
Forking After Creating the Client
Python servers like Gunicorn and uWSGI fork worker processes. A PyMongo client created in the parent before the fork is copied into every child, and its pool and monitoring threads don't survive forking correctly. PyMongo warns about this. Create the client lazily in each worker (or in a post-fork hook) instead. Node.js cluster mode and PM2 don't have this problem because each worker runs your module from scratch, but they still multiply your pool count by the number of workers.
Serverless Scale-Out
Serverless platforms can run hundreds or thousands of concurrent instances, each with its own client. This is a big enough topic that it gets its own post: Using MongoDB in Serverless Functions Without Exhausting Connections.
Slow Queries Backing Up the Pool
Sometimes the client setup is correct and the pool is simply saturated because operations take too long. If a query that used to take 5 ms now takes 2 seconds because it's doing a collection scan, each connection is busy 400 times longer, and every pool in your fleet grows to its maximum. Raising maxPoolSize here makes things worse: more concurrent slow queries compete for the same CPU and disk. Fix the query instead. Using explain() to Analyze and Debug Slow MongoDB Queries walks through how.
Best Practices
Create exactly one client per process and share it. A client is thread-safe, goroutine-safe, and async-safe in every official driver. Multiple clients in one process means multiple pools for no benefit.
Set appName on every client. It turns "who is holding 1,200 connections?" from a mystery into a single aggregation.
Size pools for your process count. Divide the connection budget you're willing to spend by the number of processes, and set maxPoolSize accordingly. Leave headroom for deploys, where old and new versions run side by side.
Set a wait queue timeout. A bounded wait (a few seconds) makes pool exhaustion fail loudly and quickly instead of hanging requests until the load balancer gives up.
Use maxIdleTimeMS for bursty traffic. It lets pools shrink after spikes, so idle connections don't hold server memory.
Close clients on shutdown. Call client.close() in your shutdown handler so rolling deploys release connections promptly instead of waiting for the server to notice dead sockets.
Treat connection count as a first-class metric. Alert on it well below your cluster limit, and watch totalCreated as a rate. A steadily climbing creation rate is an early warning for a client-per-request bug.
Conclusion
Too many connections is rarely a database problem. It's a multiplication problem: processes times clients times pool size. Once you create one client per process, set appName, size pools for your fleet rather than for a single server, and use $currentOp to see who's holding what, connection limits stop being something you hit by surprise.
Your next step: run the $currentOp grouping query above against your production cluster today. If any service's count is far higher than its process count times maxPoolSize, you've found a client being created somewhere it shouldn't be.


