Type something to search...
Multi-Tenant Application Design Patterns in MongoDB

Multi-Tenant Application Design Patterns in MongoDB

Almost every SaaS product ends up with the same question in its first year: how do you keep each customer's data separate while running a single application? At ten customers, the answer barely matters. At ten thousand, the wrong choice shows up as a cluster with too many files to manage, a noisy customer slowing everyone else down, or, worst of all, one tenant seeing another tenant's invoices.

MongoDB gives you several ways to model multi-tenancy, and each one trades isolation against operational cost. Separate databases give strong boundaries and simple per-customer backups. A shared collection with a tenantId field scales to huge numbers of tenants with almost no overhead. Most mature systems end up with a hybrid of the two.

This guide covers the three main patterns, how to compare them, how to enforce tenant isolation in application code, how indexing and sharding change when tenants share collections, and the mistakes that tend to surface only after you've grown.

The Three Core Patterns

Every MongoDB multi-tenant design is a variation on one of three models.

Database per Tenant

Each tenant gets its own database: tenant_acme, tenant_globex, and so on. Collection names inside are identical across tenants, and your application picks the database based on who's making the request.

// mongosh
use tenant_acme
db.invoices.find({ status: "open" })

use tenant_globex
db.invoices.find({ status: "open" })

This is the most isolated option short of separate clusters. You can create a database user scoped to a single tenant database, back up or restore one tenant with mongodump --db, drop a tenant cleanly when they leave, and even tweak indexes for one large customer without touching anyone else.

The cost is overhead. Every collection and every index in MongoDB is backed by its own file in the WiredTiger storage engine. Ten collections with four indexes each is fifty files per tenant. Multiply by five thousand tenants and you're at a quarter million files, each consuming memory for metadata and handles. Startup, replication initial sync, and backups all slow down as that number grows.

Collection per Tenant

A middle ground: one database, but each tenant gets its own set of collections, like acme_invoices and globex_invoices. In practice this has most of the overhead of database-per-tenant with fewer of the benefits. You can't grant access at the database level anymore (collection-level privileges are possible but awkward), and the collection count explodes just as quickly.

It occasionally makes sense for a single large, tenant-specific dataset in an otherwise shared design. As a primary strategy, it's rarely the right pick.

Shared Collection with a Tenant Field

Every tenant's documents live in the same collections, and each document carries a tenant identifier:

db.invoices.insertOne({
  tenantId: "acme",
  number: "INV-1042",
  customer: "Wile E. Coyote",
  total: NumberDecimal("1499.00"),
  status: "open",
  createdAt: new Date(),
});

db.invoices.find({ tenantId: "acme", status: "open" });

This is sometimes called the pool model. The number of files stays constant no matter how many tenants you add, indexes are shared, and onboarding a new tenant is just inserting a row into a tenants collection. It's how most large SaaS products on MongoDB work.

The trade-off is that isolation is now entirely your application's responsibility. The database has no idea that tenantId is special. Forget the filter in one query and you've built a data leak.

Comparing the Patterns

ConcernDatabase per tenantCollection per tenantShared collection
Isolation strengthStrong (enforceable by roles)MediumApplication-enforced
Scales to many tenantsHundreds to low thousandsPoorlyMillions
Per-tenant backup and restoreEasy (mongodump --db)Possible per namespaceRequires filtered export
Tenant offboardingdropDatabase()Drop collectionsdeleteMany({ tenantId })
Per-tenant index tuningYesYesNo (indexes are shared)
Schema migrationsRun N timesRun N timesRun once
Noisy neighbor riskLower (still shared hardware)LowerHigher
Connection and pooling overheadSame client, database switchSameSame
Cross-tenant analyticsHard (N queries)HardEasy (one aggregation)

Notice that none of these models gives you separate hardware. If you need guaranteed resources or a contractual boundary (common in regulated industries), the real fourth option is a cluster per tenant, usually reserved for your largest or most sensitive customers.

Choosing a Model

A few questions narrow the choice quickly:

  • How many tenants do you expect? If the answer is "thousands of small ones," go shared. Database-per-tenant starts to hurt well before you reach five figures.
  • Do contracts or regulations require hard separation? If a customer's security review asks for dedicated storage or their own encryption keys, the pool model won't satisfy them, and you'll want a database or cluster for that tenant.
  • Are tenants roughly the same size? A shared collection works best when tenants are small relative to the whole. One tenant with 60% of the data behaves very differently from ten thousand tenants with 0.01% each.
  • Will you need per-tenant restores? "Restore Acme to yesterday at 3pm without touching anyone else" is trivial with separate databases and genuinely hard with shared collections.

For most B2B SaaS products, the pragmatic answer is: start with a shared collection, and design your data access layer so a tenant can later be moved to its own database or cluster without changing business logic.

Enforcing Isolation in the Shared Model

If you pick the pool model, the single most important thing you can do is make it impossible to query without a tenant filter. Don't rely on every developer remembering to add tenantId to every query. Centralize it.

A Tenant-Scoped Repository

Here's a small wrapper for the Node.js driver that injects the tenant on every read and write:

// tenantCollection.js
export function tenantCollection(db, name, tenantId) {
  if (!tenantId) {
    throw new Error(`tenantId is required for ${name}`);
  }

  const coll = db.collection(name);
  const scope = (filter = {}) => ({ ...filter, tenantId });

  return {
    find: (filter, options) => coll.find(scope(filter), options),
    findOne: (filter, options) => coll.findOne(scope(filter), options),
    countDocuments: (filter, options) =>
      coll.countDocuments(scope(filter), options),
    insertOne: (doc, options) => coll.insertOne({ ...doc, tenantId }, options),
    insertMany: (docs, options) =>
      coll.insertMany(
        docs.map((d) => ({ ...d, tenantId })),
        options,
      ),
    updateOne: (filter, update, options) =>
      coll.updateOne(scope(filter), update, options),
    updateMany: (filter, update, options) =>
      coll.updateMany(scope(filter), update, options),
    deleteOne: (filter, options) => coll.deleteOne(scope(filter), options),
    deleteMany: (filter, options) => coll.deleteMany(scope(filter), options),
    aggregate: (pipeline, options) =>
      coll.aggregate([{ $match: { tenantId } }, ...pipeline], options),
  };
}

Two details matter here. The spread puts tenantId last, so a caller can't override it by passing their own tenantId in the filter. And the aggregation wrapper prepends a $match, so every pipeline starts scoped to the tenant (and can use a tenant-prefixed index).

In your request handling, derive the tenant from something trustworthy, like a verified JWT claim or the session, never from a query string or request body:

app.get("/api/invoices", async (req, res) => {
  const tenantId = req.auth.tenantId; // set by auth middleware
  const invoices = tenantCollection(db, "invoices", tenantId);

  const open = await invoices
    .find({ status: "open" })
    .sort({ createdAt: -1 })
    .limit(50)
    .toArray();

  res.json(open);
});

Watch the Escape Hatches

A wrapper protects the obvious paths. The subtle leaks come from operations that reach into other collections:

  • $lookup and $graphLookup join against another collection without your tenant filter. Use the pipeline form of $lookup and match on tenantId inside it.
  • $unionWith pulls documents from a second collection, again unscoped.
  • Change streams see every tenant's changes unless you add a $match stage on fullDocument.tenantId.
  • Atlas Search and vector search indexes span the whole collection, so include tenantId as a filter field in the index definition and in every query.

Here's a scoped $lookup:

db.invoices.aggregate([
  { $match: { tenantId: "acme", status: "open" } },
  {
    $lookup: {
      from: "customers",
      let: { customerId: "$customerId", tenantId: "$tenantId" },
      pipeline: [
        {
          $match: {
            $expr: {
              $and: [
                { $eq: ["$_id", "$$customerId"] },
                { $eq: ["$tenantId", "$$tenantId"] },
              ],
            },
          },
        },
      ],
      as: "customer",
    },
  },
]);

Add Schema Validation as a Backstop

You can't make MongoDB enforce "every query includes tenantId," but you can at least guarantee that every document has one. A $jsonSchema validator catches writes from scripts or services that bypass your repository layer:

db.runCommand({
  collMod: "invoices",
  validator: {
    $jsonSchema: {
      bsonType: "object",
      required: ["tenantId"],
      properties: {
        tenantId: { bsonType: "string", minLength: 1 },
      },
    },
  },
  validationAction: "error",
});

See Schema Validation in MongoDB with JSON Schema for the full range of rules you can apply.

Indexing for Shared Collections

In the pool model, tenantId should be the first field of almost every index. Every query filters on it, so an index that starts with it lets MongoDB jump straight to one tenant's slice of the data.

db.invoices.createIndex({ tenantId: 1, status: 1, createdAt: -1 });
db.invoices.createIndex({ tenantId: 1, number: 1 }, { unique: true });
db.customers.createIndex({ tenantId: 1, email: 1 }, { unique: true });

The unique indexes are worth a second look. Invoice numbers and customer emails are usually unique per tenant, not globally. Two companies can both have an INV-1042. Prefixing the unique index with tenantId expresses exactly that rule. A plain { email: 1 } unique index would stop two tenants from having customers with the same email, which is a bug that tends to be discovered by an angry support ticket.

An index like { status: 1, createdAt: -1 } without the tenant prefix is almost always a mistake in this model. MongoDB would scan every tenant's open invoices and then filter, which gets slower as your whole customer base grows rather than as the individual tenant grows.

Scaling with Sharding

When a shared collection outgrows a single replica set, tenantId becomes a natural part of the shard key. A compound key keeps each tenant's data together while still allowing large tenants to split across chunks:

sh.shardCollection("app.invoices", { tenantId: 1, _id: 1 });

Queries that include tenantId are routed to only the shards holding that tenant's range, instead of broadcasting to the whole cluster. Because _id follows, one giant tenant can still be split into many chunks and balanced across shards.

A pure { tenantId: 1 } key is risky: all of one tenant's documents share a single shard key value, so they can never be split, and a big tenant becomes a jumbo chunk. A hashed tenantId spreads tenants evenly but has the same problem for any single large tenant. The deeper trade-offs are covered in How to Choose a Good Shard Key in MongoDB.

Zone Sharding for Data Residency

Sharding also solves a common enterprise requirement: keeping a tenant's data in a specific region. With zone sharding, you pin ranges of the shard key to shards in particular locations:

sh.addShardToZone("shard-eu-1", "EU");
sh.addShardToZone("shard-us-1", "US");

sh.updateZoneKeyRange(
  "app.invoices",
  { tenantId: "eu_", _id: MinKey },
  { tenantId: "eu`", _id: MinKey },
  "EU",
);

If your EU tenants all have IDs starting with eu_, this range routes them to EU shards (the backtick is the character immediately after the underscore in ASCII, so the range covers every eu_ prefix). Designing tenant IDs with a region prefix up front makes this much easier than retrofitting it later.

The Hybrid Model

Real systems rarely stay pure. A common evolution looks like this:

  1. Small and medium tenants live in shared collections.
  2. Large or regulated tenants get their own database, or their own cluster.
  3. A tenant catalog records where each tenant lives.
db.tenants.insertMany([
  { _id: "acme", plan: "pro", placement: { type: "shared" } },
  {
    _id: "globex",
    plan: "enterprise",
    placement: {
      type: "dedicated",
      uri: "mongodb+srv://globex-cluster.example.net",
    },
  },
]);

Your data access layer reads the catalog (cached in memory), then hands back either a scoped wrapper over the shared database or a plain collection from the tenant's dedicated database. Business logic calls the same invoices.find() either way.

const clients = new Map();

async function getTenantCollection(tenant, name) {
  if (tenant.placement.type === "shared") {
    return tenantCollection(sharedDb, name, tenant._id);
  }

  let client = clients.get(tenant._id);
  if (!client) {
    client = new MongoClient(tenant.placement.uri, { maxPoolSize: 20 });
    clients.set(tenant._id, client);
  }
  // Keep tenantId on documents even in dedicated DBs, so data can move back and forth
  return tenantCollection(client.db("app"), name, tenant._id);
}

Keeping tenantId on every document even in dedicated databases is a small cost that pays off enormously: moving a tenant between placements becomes a copy-and-switch, not a transformation.

Be careful with the client cache. Each MongoClient maintains its own connection pool per server, so a hundred dedicated clusters times several app instances can add up to a lot of connections. That's a good reason to keep dedicated placements for the tenants who really need them.

Handling the Noisy Neighbor

In any shared model, one tenant running a heavy export can slow down everyone else. Some defenses:

  • Rate limit per tenant at the API layer, not just per user.
  • Cap expensive queries with maxTimeMS so a runaway aggregation gets killed instead of saturating the cluster.
  • Route analytics to secondaries with a secondaryPreferred read preference or dedicated analytics nodes, keeping reporting off the primary.
  • Paginate everything, with range-based pagination on indexed fields rather than large skip values.
  • Promote outliers. When one tenant's workload is consistently out of proportion, move them to dedicated placement. The hybrid catalog makes this a routine operation.
const report = await invoices
  .aggregate([{ $group: { _id: "$status", total: { $sum: "$total" } } }], {
    maxTimeMS: 5000,
    readPreference: "secondaryPreferred",
  })
  .toArray();

Security Considerations

In the database-per-tenant model, create a MongoDB user per tenant (or per tenant-facing service) with the readWrite role on only that tenant's database. Even if application code has a bug, the connection simply can't read another tenant's database.

use admin
db.createUser({
  user: "svc_acme",
  pwd: passwordPrompt(),
  roles: [{ role: "readWrite", db: "tenant_acme" }],
});

In practice, most applications still connect with a single service account for simplicity, which means the isolation benefit only materializes if you actually use per-tenant credentials.

In the shared model, defense in depth comes from the repository wrapper, schema validation, code review rules (lint for raw db.collection() calls outside the data layer), and automated tests that assert one tenant can't read another's records. Write at least one integration test per collection that inserts data for tenant A and verifies that tenant B's scoped queries return nothing.

For tenants who require their own encryption keys, look at Client-Side Field Level Encryption or Queryable Encryption, where each tenant's sensitive fields can be encrypted with a key only their context can access.

Common Pitfalls

Taking the tenant ID from user input. If the API accepts ?tenantId=acme, anyone can change it. Derive the tenant from authenticated context only.

Global unique indexes. A { email: 1 } unique index across a shared collection prevents two tenants from having the same user email. Prefix unique indexes with tenantId.

Starting with database-per-tenant for a self-serve product. Free trials and small accounts pile up fast. If anyone can sign up with a credit card, you'll hit collection-count pressure long before you hit revenue that justifies it.

Forgetting background jobs. Cron jobs, queues, and migration scripts often bypass the request layer and its tenant scoping. Make them use the same data access layer, with an explicit tenant loop.

Unscoped joins and change streams. $lookup, $unionWith, and change streams are the most common places a tenant filter goes missing, because they don't pass through a simple find() wrapper.

Not planning for offboarding. Deleting a tenant from shared collections means deleteMany across every collection, plus handling backups that still contain their data. Keep a registry of every collection that stores tenant data so offboarding is a script, not an investigation.

Conclusion

Multi-tenancy in MongoDB comes down to where you put the boundary. Database-per-tenant gives you strong, role-enforced isolation and easy per-tenant operations, but it doesn't scale to large numbers of tenants. A shared collection with a tenantId scales almost indefinitely, provided you enforce the filter centrally, prefix your indexes with the tenant, and include it in your shard key. The hybrid model lets you use both, keeping small tenants pooled and moving large or regulated ones to dedicated placement.

If you're starting today, pick one collection in your app and route every query against it through a tenant-scoped wrapper like the one above. Once that pattern exists, extending it to the rest of your data layer, and later to dedicated placements, is straightforward.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading