Type something to search...
Using MongoDB in Serverless Functions Without Exhausting Connections

Using MongoDB in Serverless Functions Without Exhausting Connections

Serverless functions and traditional databases have a fundamental mismatch. A function platform happily spins up hundreds or thousands of instances when traffic spikes, and each instance is a separate process with its own memory. MongoDB drivers are designed around long-lived processes that open a connection pool once and reuse it for hours. Put the two together naively and every new instance opens its own pool, your connection count climbs with your traffic, and eventually the database starts refusing connections.

The good news is that the fix is well understood. Most of it comes down to three things: create the client outside the handler so warm invocations reuse it, keep each instance's pool tiny, and put a ceiling on how many instances can run at once. A few platform-specific details (like Lambda's event loop behavior and Vercel's Fluid compute) fill in the rest.

This guide covers why serverless functions exhaust connections, the client reuse pattern for AWS Lambda, Vercel, Netlify, and Google Cloud Run functions, how to size pools and concurrency, networking and authentication, and the options that don't work (or no longer exist).

Why Serverless Breaks Connection Math

In a regular server, connection count is roughly processes x maxPoolSize, and the number of processes is something you control. In serverless, the platform controls it. Each concurrent request can land on its own instance, and the platform creates instances as fast as traffic demands.

Consider what happens with the default driver settings:

  • A traffic spike triggers 400 concurrent function instances.
  • Each creates a MongoClient with the default maxPoolSize of 100.
  • Each client also opens monitoring connections to every replica set member.

Even if each instance only uses one or two connections at a time, that's hundreds of application connections plus monitoring connections to all three members. If the handler creates a new client on every invocation instead of reusing one, it's far worse: connections are opened and abandoned continuously, and the server spends its CPU on TLS handshakes and authentication instead of queries.

The free Atlas M0 tier allows 500 connections, and Flex clusters have their own limits too. Dedicated tiers allow more, but even there, an unbounded fan-out will find the ceiling. (For the general theory of pools and limits, see MongoDB Connection Pooling: Avoiding Too Many Connections.)

The Core Pattern: Create the Client Outside the Handler

Serverless platforms reuse an instance for many invocations when traffic is steady. Anything you create at module scope survives between invocations on the same instance. So the single most important change is moving client creation out of the handler.

Here's the wrong way:

// Anti-pattern: a new client, pool, and TLS handshake on every invocation
export const handler = async (event) => {
  const client = new MongoClient(process.env.MONGODB_URI);
  const orders = client.db("shop").collection("orders");
  const order = await orders.findOne({ orderId: event.pathParameters.id });
  return { statusCode: 200, body: JSON.stringify(order) };
};

And the right way:

// lambda/getOrder.mjs
import { MongoClient } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI, {
  appName: "orders-lambda",
  maxPoolSize: 5,
  minPoolSize: 0,
  maxIdleTimeMS: 60_000,
  serverSelectionTimeoutMS: 5_000,
});

// Start connecting during the init phase, which runs once per instance
const clientPromise = client.connect();

export const handler = async (event) => {
  await clientPromise;
  const orders = client.db("shop").collection("orders");

  const order = await orders.findOne(
    { orderId: event.pathParameters.id },
    { projection: { _id: 0 } },
  );

  if (!order) {
    return { statusCode: 404, body: JSON.stringify({ error: "Not found" }) };
  }
  return { statusCode: 200, body: JSON.stringify(order) };
};

The first invocation on a new instance (a cold start) pays for the connection. Every warm invocation after that reuses the pool and runs the query immediately. Kicking off connect() at module scope also lets the handshake overlap with the rest of the init phase.

Note what this code doesn't do: it never calls client.close() in the handler. Closing at the end of each invocation would throw away the pool you're trying to reuse. The platform freezes and eventually recycles the instance, and the server cleans up the idle connections.

Lambda and callbackWaitsForEmptyEventLoop

With older callback-style Lambda handlers, Lambda waits for the Node.js event loop to be empty before returning the response. An open MongoDB client keeps timers and sockets on the event loop, so the invocation would hang until it times out. The fix was:

export const handler = (event, context, callback) => {
  context.callbackWaitsForEmptyEventLoop = false;
  // ...
};

With async handlers like the example above, Lambda returns as soon as the promise resolves, so you don't need this setting. If you're maintaining callback-style handlers, set it at the top of the handler.

Size Pools for One Request at a Time

Most function platforms send one request at a time to each instance. That means an instance rarely needs more than one or two connections at once. A maxPoolSize between 1 and 10 is plenty; 5 is a reasonable default that leaves room for parallel queries (Promise.all over a few lookups) within a single invocation.

The math changes dramatically:

SetupInstancesPool per instanceWorst-case app connections
Default settings40010040,000
Tuned40052,000
Tuned plus concurrency cap of 1001005500

maxIdleTimeMS helps too. After a spike, instances that are still warm but idle close their unused connections instead of holding them until the platform recycles the instance.

Cap Concurrency

Pool tuning reduces connections per instance, but only a concurrency limit bounds the number of instances. Every major platform has a setting for it:

  • AWS Lambda: reserved concurrency on the function (for example, 100). Requests beyond the cap are throttled, and API Gateway returns a 429 or SQS retries the message later.
  • Google Cloud Run functions: maximum instances setting, plus a per-instance concurrency setting (Cloud Run instances can serve several requests at once, in which case give the pool a few more connections).
  • Vercel and Netlify: concurrency is managed by the platform and your plan; check your plan's limits and scaling behavior.

Pick the cap by working backwards from your connection budget:

max instances = (connection budget for this function) / maxPoolSize

If you have several functions talking to the same cluster, split the budget between them. A throttled function is a far better failure mode than a database that rejects connections for every service at once.

For write-heavy background work, consider putting a queue (SQS, Pub/Sub) in front of the function and limiting how many consumers run concurrently. The queue absorbs the burst, and the database sees a steady, bounded load.

Platform-Specific Patterns

Vercel and Next.js

On Vercel, the same module-scope rule applies. For Next.js route handlers and Server Actions, a shared module with a cached client works in both production and development:

// lib/mongodb.ts
import { MongoClient } from "mongodb";

const uri = process.env.MONGODB_URI!;
const options = {
  appName: "storefront",
  maxPoolSize: 5,
  maxIdleTimeMS: 60_000,
};

const globalForMongo = globalThis as unknown as { _mongoClient?: MongoClient };

export const client =
  globalForMongo._mongoClient ?? new MongoClient(uri, options);

if (process.env.NODE_ENV !== "production") {
  // Survive hot reloads in development without leaking clients
  globalForMongo._mongoClient = client;
}

Vercel's Fluid compute lets a single instance handle multiple concurrent invocations and keeps instances alive longer, which makes pooling more effective. Vercel also provides an attachDatabasePool helper in the @vercel/functions package that helps the platform close idle pool connections before an instance is suspended:

import { attachDatabasePool } from "@vercel/functions";
import { client } from "@/lib/mongodb";

attachDatabasePool(client);

Check Vercel's current documentation for how this interacts with your runtime and plan. For more on this stack specifically, see Using MongoDB with Next.js App Router and Server Actions.

One important limit: Edge runtimes can't run the MongoDB Node.js driver. The driver needs raw TCP sockets and Node.js APIs. Run database code in the Node.js runtime (the default for route handlers), not in runtime = "edge" functions or middleware.

Python on Lambda

The same pattern applies in Python. Module-level code runs once per instance:

# handler.py
import json
import os

from pymongo import MongoClient

client = MongoClient(
    os.environ["MONGODB_URI"],
    appname="orders-lambda-py",
    maxPoolSize=5,
    maxIdleTimeMS=60000,
    serverSelectionTimeoutMS=5000,
)
orders = client["shop"]["orders"]


def handler(event, context):
    order = orders.find_one(
        {"orderId": event["pathParameters"]["id"]},
        {"_id": 0},
    )
    if order is None:
        return {"statusCode": 404, "body": json.dumps({"error": "Not found"})}
    return {"statusCode": 200, "body": json.dumps(order, default=str)}

PyMongo connects lazily on the first operation, so there's no explicit connect call. Since Lambda runs one request per instance, the synchronous client is the right choice here.

Networking and Authentication

Connections aren't just a count problem; each new one has to be allowed in and authenticated.

IP access lists. Lambda functions outside a VPC use a pool of AWS egress IPs that change, so you can't reliably allowlist them. Opening Atlas to 0.0.0.0/0 works but depends entirely on credentials for security. Better options are running the function in a VPC with a NAT gateway (giving it a fixed egress IP) or using a private endpoint (AWS PrivateLink) or VPC peering to your Atlas cluster, available on dedicated tiers.

Authentication without stored passwords. On AWS, Atlas supports AWS IAM authentication (authMechanism=MONGODB-AWS). The function's execution role is mapped to a database user in Atlas, and the driver picks up the role's temporary credentials from the environment automatically:

mongodb+srv://cluster0.example.mongodb.net/shop?authSource=%24external&authMechanism=MONGODB-AWS&appName=orders-lambda

No password lives in your environment variables, and rotating credentials is handled by AWS. In the Node.js driver, this requires the optional @aws-sdk/credential-providers package to be installed alongside mongodb.

Cold start cost. Each new connection does DNS SRV resolution, a TCP connect, a TLS handshake, and authentication, which can add a few hundred milliseconds to a cold start. Deploying functions in the same region as your cluster is the biggest single improvement. Provisioned concurrency (Lambda) or minimum instances (Cloud Run) keeps a few instances warm if cold start latency matters.

What Not to Use

A few approaches show up in older articles but aren't good choices today:

  • The Atlas Data API and Atlas HTTPS Endpoints. These let functions query Atlas over HTTP without a driver, which neatly sidestepped connection limits. Both were deprecated and reached end of life in September 2025, so don't build on them. If you need an HTTP layer, run your own small API service (with a normal long-lived pool) and have functions call it.
  • Serverless instances. Atlas replaced its old Serverless instances (and the M2/M5 shared tiers) with Flex clusters. Flex is a good fit for spiky, low-to-moderate traffic, but it's a cluster tier choice, not a connection-management feature. The same client reuse and concurrency rules apply.
  • Closing the client after every invocation. It avoids leaked connections but pays the full connection cost on every request and creates constant churn on the server. Reuse is better in every respect.

Common Mistakes

Creating the client inside the handler. This is the root cause of most serverless connection problems. Module scope, always.

Leaving maxPoolSize at 100. A function instance handling one request at a time will never use 100 connections, but the ceiling still applies during slow-query backups. Set it to single digits.

No concurrency limit. Without a cap, your connection count is set by your traffic, which means a viral moment or a retry storm can take down the database for every service sharing it.

Running driver code in the Edge runtime. It fails at build or run time. Keep MongoDB access in Node.js runtime functions.

Forgetting appName. When 600 connections show up during a spike, appName is what tells you they belong to orders-lambda and not to the API service.

Assuming warm instances stay warm. Platforms recycle instances unpredictably. Your code must work correctly on a cold start every time; reuse is an optimization, not a guarantee.

Conclusion

MongoDB works well in serverless functions once you stop fighting the platform's scaling model. Create the client at module scope so warm invocations reuse it, keep pools tiny because each instance handles about one request at a time, and cap concurrency so the number of instances can't outgrow your cluster's connection limit. Add IAM authentication and private networking where your platform supports them, and skip the deprecated HTTP APIs.

Your next step: open your busiest function, move its MongoClient to module scope if it isn't there already, set maxPoolSize: 5 and an appName, and set a reserved concurrency limit that fits your cluster's connection budget. Then watch the connection graph during your next traffic spike.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading