Type something to search...
Capped Collections in MongoDB: Use Cases and Limitations

Capped Collections in MongoDB: Use Cases and Limitations

Some data only matters while it's recent. The last few thousand log lines from a service, the latest status messages from a job runner, a stream of events you want to tail in real time. You don't care about last month's entries, and you definitely don't want the collection to grow until it eats your disk.

A capped collection is MongoDB's answer to that. It has a fixed maximum size, set when you create it. Documents are stored in insertion order, and once the collection is full, each new insert overwrites the oldest document. Think of it as a ring buffer built into the database. There's no cleanup job, no TTL monitor, and no growth: the collection simply never gets bigger than you allowed.

This guide covers creating and sizing capped collections, reading them in natural order, tailing them with tailable cursors, resizing and converting them, and the real limitations that make capped collections the wrong choice more often than people expect.

Creating a Capped Collection

Capped collections must be created explicitly with capped: true and a size in bytes:

db.createCollection("serviceLog", {
  capped: true,
  size: 64 * 1024 * 1024, // 64 MB
});

You can also set a maximum number of documents with max:

db.createCollection("recentAlerts", {
  capped: true,
  size: 10 * 1024 * 1024, // 10 MB
  max: 5000,
});

A few details about these limits:

  • size is always required, even when you set max. It's the hard limit in bytes.
  • Whichever limit is reached first wins. If 5,000 alerts take only 2 MB, max controls. If each alert is large and 10 MB fills up at 3,000 documents, size controls.
  • Size is rounded. MongoDB rounds size up to a multiple of 256 bytes, and very small values are raised to a minimum of 4,096 bytes.

You can confirm a collection is capped and see its limits:

db.recentAlerts.isCapped();
// true

const s = db.recentAlerts.stats();
({
  capped: s.capped,
  maxSize: s.maxSize,
  max: s.max,
  count: s.count,
  size: s.size,
});
{ capped: true, maxSize: 10485760, max: 5000, count: 0, size: 0 }

From Node.js, it's the same options on createCollection:

import { MongoClient } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
const db = client.db("ops");

const existing = await db.listCollections({ name: "serviceLog" }).toArray();
if (existing.length === 0) {
  await db.createCollection("serviceLog", {
    capped: true,
    size: 64 * 1024 * 1024,
  });
}

The existence check matters. If your code simply inserts into a collection that doesn't exist yet, MongoDB creates a normal, uncapped collection, and you won't notice until it has grown for months.

Insertion Order and Natural Order

In a capped collection, documents are guaranteed to be stored and returned in insertion order when you read in natural order, meaning with no sort:

db.serviceLog.insertMany([
  { level: "info", msg: "worker started", at: new Date() },
  { level: "warn", msg: "queue depth 950", at: new Date() },
  { level: "error", msg: "upstream timeout", at: new Date() },
]);

db.serviceLog.find(); // oldest first
db.serviceLog.find().sort({ $natural: -1 }).limit(10); // newest ten

Sorting by $natural: -1 walks the collection backwards from the most recent insert, which makes "show me the last N entries" very cheap. That's one of the nicest properties of capped collections: the most common log query needs no index at all.

Because order is fixed by insertion, a capped collection doesn't need an index to keep "oldest" and "newest" straight. It still has an _id index by default, and you can add secondary indexes if you need to query by other fields:

db.serviceLog.createIndex({ level: 1 });
db.serviceLog.find({ level: "error" }).sort({ $natural: -1 }).limit(20);

Keep in mind that each index adds overhead to every insert, and capped collections are usually chosen for high-throughput, low-overhead writes.

Tailable Cursors: Following a Collection Like tail -f

The feature that makes capped collections genuinely unique is the tailable cursor. A normal cursor closes when it reaches the end of the results. A tailable cursor stays open, and when new documents are inserted, it returns them. It's the database equivalent of tail -f.

In mongosh:

const cursor = db.serviceLog.find().tailable({ awaitData: true });
while (!cursor.isClosed()) {
  if (cursor.hasNext()) printjson(cursor.next());
}

With awaitData, the server waits briefly for new data before returning an empty batch, so the client isn't hammering it in a busy loop.

In Node.js, a tailable cursor works nicely with for await:

import { MongoClient } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
const log = client.db("ops").collection("serviceLog");

async function follow(onEntry) {
  let lastId = null;

  while (true) {
    const filter = lastId ? { _id: { $gt: lastId } } : {};
    const cursor = log.find(filter, {
      tailable: true,
      awaitData: true,
      maxAwaitTimeMS: 2000,
    });

    try {
      for await (const doc of cursor) {
        lastId = doc._id;
        await onEntry(doc);
      }
    } catch (err) {
      console.warn("tailable cursor ended:", err.message);
    }

    // Cursor died (empty collection, or we fell behind). Back off and reopen.
    await new Promise((r) => setTimeout(r, 1000));
  }
}

follow((doc) => console.log(`[${doc.level}] ${doc.msg}`));

The reconnect loop is not optional. A tailable cursor can die for several reasons:

  • The collection is empty when the cursor is opened. There's no position to wait at, so the cursor is dead immediately.
  • The reader falls behind. If writers overwrite the document the cursor was positioned at, the cursor is lost.
  • Network errors or failovers close cursors like any other.

Resuming from the last seen _id works because default ObjectIds increase roughly with time, but the filter on _id has to be evaluated by scanning, since tailable cursors don't use indexes to find their starting position. For small capped collections that's fine. If you need a robust, resumable feed of changes, change streams are usually the better tool.

In Python with PyMongo:

import time
from pymongo import MongoClient, CursorType

log = MongoClient("mongodb://localhost:27017")["ops"]["serviceLog"]

while True:
    cursor = log.find(cursor_type=CursorType.TAILABLE_AWAIT)
    while cursor.alive:
        for doc in cursor:
            print(f"[{doc['level']}] {doc['msg']}")
        time.sleep(0.5)
    time.sleep(1)

Good Use Cases

Capped collections fit a specific shape of problem: high-volume, append-only data where only the most recent window matters and a fixed storage budget is more important than a fixed time window.

Recent application logs. Keep the last 100 MB of logs from a service in the database for quick debugging, with a real log pipeline handling long-term retention.

Rolling status and activity feeds. "Recent activity" panels, job runner output, deployment logs streamed to a dashboard.

Lightweight message relays. A small capped collection plus tailable cursors can act as a simple pub/sub channel between processes that already share a MongoDB connection. It's not a replacement for a message broker, but for low-stakes internal notifications it's surprisingly handy.

Caches with a size budget. When you'd rather keep "the most recent N megabytes" than "everything from the last N hours."

MongoDB itself uses the same idea internally. The replication oplog (local.oplog.rs) behaves like a capped collection, retaining a bounded window of operations. The oplog explainer goes into how that works.

The Limitations

This is where capped collections lose most of the arguments. The restrictions are real, and several of them are dealbreakers for general-purpose data.

You Can't Delete Individual Documents

Deleting specific documents from a capped collection isn't allowed:

db.serviceLog.deleteOne({ level: "debug" });
MongoServerError: cannot remove from a capped collection: ops.serviceLog

Documents only leave when they're overwritten by new inserts. If you need to remove data on request (a user asks you to delete their information, for instance), a capped collection makes that impossible without dropping the whole collection. That alone rules capped collections out for anything containing personal data you might need to erase.

Updates Can't Change Document Size

You can update documents, but an update that changes a document's size fails. Setting a status field from "ok" to "failed" changes the size; so does pushing to an array. In practice this means you should treat capped collections as insert-only.

No Sharding, No Transactions

Capped collections can't be sharded, so they're limited to what a single replica set can handle. You also can't write to a capped collection inside a multi-document transaction.

Size-Based, Not Time-Based

The collection keeps as much data as fits. During a traffic spike or a burst of error logs, the window of retained data shrinks from days to minutes, exactly when you'd most want the history. If you need a guarantee like "keep at least 7 days," a capped collection can't give it to you.

TTL Indexes Don't Apply

You can't add a TTL index to a capped collection. Expiry is purely by space.

Resizing and Converting

In MongoDB 6.0 and later, you can change a capped collection's limits without recreating it using collMod:

db.runCommand({
  collMod: "serviceLog",
  cappedSize: 256 * 1024 * 1024,
  cappedMax: 200000,
});

Shrinking the limits causes the oldest documents to be removed until the collection fits.

To turn an existing regular collection into a capped one, use convertToCapped:

db.runCommand({ convertToCapped: "events", size: 100 * 1024 * 1024 });

This rebuilds the collection and holds an exclusive lock on the database while it runs, so don't run it on a large, busy production collection without planning for the downtime. It also doesn't copy secondary indexes, only _id, so recreate any others afterward.

There's no command to convert a capped collection back into a regular one. You'd create a new collection and copy the data over, for example with an aggregation ending in $out.

Capped Collections vs. the Alternatives

For most "old data should go away" problems, you have better options than a capped collection:

NeedBest tool
Keep data for a fixed time, then deleteTTL index
High-volume measurements with time-based expiryTime series collection
Real-time feed of changes, resumable after crashesChange streams
Hard cap on storage, insertion order, tail -fCapped collection
Ability to delete individual recordsRegular collection (anything but capped)

MongoDB's own documentation recommends TTL indexes over capped collections in many cases, because they give you more flexibility and work with sharding. TTL indexes handle time-based retention, and time series collections handle high-volume metrics with efficient bucket-level expiry.

Where capped collections still win is when the storage cap is the requirement. If you're running MongoDB on a small edge device or a shared development server, and the rule is "this log must never exceed 50 MB no matter what," nothing else guarantees that as simply.

A Practical Example: A Bounded Debug Log

Here's a small logger that writes structured entries to a capped collection and exposes the most recent ones:

import { MongoClient } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
const db = client.db("ops");

async function ensureLog(name, bytes) {
  const [info] = await db.listCollections({ name }).toArray();
  if (!info) {
    await db.createCollection(name, { capped: true, size: bytes });
  } else if (!info.options?.capped) {
    throw new Error(`${name} exists but is not capped`);
  }
  return db.collection(name);
}

const log = await ensureLog("debugLog", 32 * 1024 * 1024);

export function write(level, msg, meta = {}) {
  // Fire-and-forget with w: 0 is tempting for logs, but acknowledged writes
  // surface misconfiguration errors. Keep the default unless you've measured.
  return log.insertOne({ at: new Date(), level, msg, ...meta });
}

export function recent(limit = 50) {
  return log.find().sort({ $natural: -1 }).limit(limit).toArray();
}

The ensureLog check guards against the most common capped collection bug: code accidentally creating a regular collection with the same name first.

Common Pitfalls

Implicitly creating the collection. Inserting into a nonexistent collection creates an uncapped one. Create capped collections explicitly at startup or in a migration, and verify with isCapped().

Undersizing. A cap that seemed generous in testing may hold only minutes of data in production. Measure average document size and insert rate, then size for the window you actually need.

Expecting to delete or edit entries. Deletes are rejected and size-changing updates fail. If you need either, use a regular collection with a TTL index.

Not handling dead tailable cursors. Tailable cursors die on empty collections and when readers fall behind. Always wrap them in a reconnect loop.

Storing personal data. Since you can't delete individual documents, you can't honor erasure requests without dropping the collection.

Running convertToCapped on a live collection. It takes an exclusive lock. Schedule it, or create a new capped collection and switch writers over.

Conclusion

Capped collections are fixed-size, insertion-ordered ring buffers. They're cheap to write, trivially easy to read newest-first, and uniquely support tailable cursors for following new data as it arrives. In exchange, you give up individual deletes, size-changing updates, sharding, transactional writes, and any guarantee about how much time the collection covers.

Use them when a hard storage cap is the actual requirement, and reach for TTL indexes or time series collections when your requirement is really about time. As a quick exercise, create a small capped collection in mongosh, open a tailable cursor in one terminal, and insert documents from another. Watching the entries stream in makes it clear exactly when this tool is the right fit.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading