Type something to search...
Tracking Document History and Versioning Changes in MongoDB

Tracking Document History and Versioning Changes in MongoDB

When you run updateOne in MongoDB, the previous value is gone. That's fine for a page view counter, but it's a real problem for a contract, a product price, a CMS article, or a user's permissions. Sooner or later someone asks "who changed this, and what did it say before?", and if you didn't plan for it, the honest answer is "we don't know."

MongoDB doesn't have built-in temporal tables, but its document model makes history surprisingly easy to add. A revision is just a document. You can store it next to the current version, in a separate collection, or as a diff, and you can write it atomically with the change itself. The hard part isn't the mechanics; it's choosing the pattern that fits how you'll actually read the history later.

This guide covers the main versioning patterns (embedded revisions, a separate history collection, and diff-based change logs), how to write them safely with transactions and optimistic concurrency, how to capture history with change streams, and how to query and restore past versions.

Decide What You Need First

Before writing code, answer a few questions, because they point to different designs:

  • Do you need full snapshots or just "what changed"? Snapshots make restoring trivial. Diffs are smaller and make audit views easier.
  • How often do you read old versions? A CMS showing "revision history" on every edit page needs fast access. A compliance archive read twice a year doesn't.
  • How many revisions per document? Five edits over a lifetime is very different from five per minute.
  • Must history be tamper-resistant? If application code can rewrite history, auditors may not trust it.
  • How long do you keep it? Forever, or 90 days?

With those answers in hand, the patterns below mostly choose themselves.

Pattern 1: Embedded Revisions

The simplest approach keeps a small array of previous versions inside the document itself:

{
  _id: ObjectId("66f0c1a2e4b0a1b2c3d4e5f6"),
  sku: "LAMP-001",
  name: "Brass Desk Lamp",
  price: NumberDecimal("89.00"),
  version: 3,
  updatedAt: ISODate("2026-09-20T10:15:00Z"),
  updatedBy: "maria",
  revisions: [
    { version: 2, price: NumberDecimal("79.00"), updatedAt: ISODate("2026-08-02T09:00:00Z"), updatedBy: "sam" },
    { version: 1, price: NumberDecimal("75.00"), updatedAt: ISODate("2026-06-14T12:30:00Z"), updatedBy: "sam" }
  ]
}

Updating pushes the old values into the array in the same atomic operation. An update pipeline can reference the current field values, so you don't need to read first:

db.products.updateOne({ sku: "LAMP-001" }, [
  {
    $set: {
      revisions: {
        $slice: [
          {
            $concatArrays: [
              [
                {
                  version: "$version",
                  price: "$price",
                  updatedAt: "$updatedAt",
                  updatedBy: "$updatedBy",
                },
              ],
              { $ifNull: ["$revisions", []] },
            ],
          },
          10,
        ],
      },
    },
  },
  {
    $set: {
      price: NumberDecimal("95.00"),
      version: { $add: ["$version", 1] },
      updatedAt: "$$NOW",
      updatedBy: "maria",
    },
  },
]);

The first stage prepends the current state to revisions and keeps only the ten most recent. The second applies the change. Because it's one update on one document, it's atomic without a transaction.

This pattern shines when history is small and you usually want it alongside the document, like "last price changes" on a product admin page. It breaks down when revisions are large or unbounded. Every revision makes the document bigger, which costs RAM and network on every read, and there's a hard 16 MB document limit. Always cap the array with $slice. If you need more than a handful of revisions, move to Pattern 2.

Pattern 2: A Separate History Collection

The most common production design keeps the current version in the main collection and copies each previous version into a history collection. This is often called the document versioning pattern.

// articles: always the latest version
{ _id: ObjectId("..."), slug: "launch-notes", title: "Launch Notes", body: "...", version: 4, updatedAt: ISODate("...") }

// articles_history: one document per past version
{
  _id: ObjectId("..."),
  articleId: ObjectId("..."),
  version: 3,
  snapshot: { title: "Launch notes", body: "...", tags: ["release"] },
  changedAt: ISODate("2026-09-18T14:02:11Z"),
  changedBy: "sam",
  reason: "Fix typo in heading"
}

Your application's normal queries never touch history, so the hot collection stays lean. History can grow as large as you like, and you can index it for the queries you need:

db.articles_history.createIndex(
  { articleId: 1, version: -1 },
  { unique: true },
);
db.articles_history.createIndex({ changedBy: 1, changedAt: -1 });

The unique index on articleId plus version doubles as a safety net: if two writers ever try to archive the same version, one fails loudly instead of creating duplicate history.

Writing Both Atomically

The main collection update and the history insert must succeed or fail together. Otherwise a crash between them leaves you with a changed document and no record of what it was. Use a transaction:

import { MongoClient, ObjectId } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
const db = client.db("cms");
const articles = db.collection("articles");
const history = db.collection("articles_history");

export async function updateArticle(
  id,
  expectedVersion,
  changes,
  user,
  reason,
) {
  const session = client.startSession();
  try {
    return await session.withTransaction(async () => {
      const current = await articles.findOne(
        { _id: new ObjectId(id), version: expectedVersion },
        { session },
      );
      if (!current) {
        throw new Error(
          "VersionConflict: article was modified by someone else",
        );
      }

      const { _id, version, updatedAt, updatedBy, ...snapshot } = current;

      await history.insertOne(
        {
          articleId: _id,
          version,
          snapshot,
          changedAt: updatedAt,
          changedBy: updatedBy,
          supersededAt: new Date(),
          supersededBy: user,
          reason,
        },
        { session },
      );

      return await articles.findOneAndUpdate(
        { _id, version: expectedVersion },
        {
          $set: { ...changes, updatedAt: new Date(), updatedBy: user },
          $inc: { version: 1 },
        },
        { session, returnDocument: "after" },
      );
    });
  } finally {
    await session.endSession();
  }
}

A few things are going on here:

  1. Optimistic concurrency. The caller passes the version it last saw. If someone else saved in the meantime, the version filter doesn't match and the update is rejected instead of silently overwriting their work. Your UI can then show "this article changed since you opened it."
  2. The snapshot excludes metadata. _id, version, and the audit fields are stored at the top level of the history document, so the snapshot contains only content.
  3. withTransaction retries transient errors automatically, including write conflicts, which is why the callback reads and writes everything through session.

Transactions require a replica set or sharded cluster. A local single-node replica set works fine for development. For more on when transactions are worth their cost, see multi-document ACID transactions.

Should History Include the Current Version?

There are two conventions. In the one above, history stores only superseded versions, and the current version lives only in the main collection. The alternative writes every version to history, including the new one, so history is a complete, self-contained log.

The "every version" approach makes queries simpler ("give me all versions" is one query) and means you can reconstruct everything from history alone. It costs one extra copy of the current state. If storage isn't tight, it's often the more convenient choice.

Pattern 3: Storing Diffs Instead of Snapshots

Full snapshots are easy to restore but wasteful for large documents with small edits. A 50 KB article where someone fixes one typo produces a 50 KB history entry. A change log stores only what changed:

{
  entityId: ObjectId("..."),
  entity: "article",
  version: 5,
  changedAt: ISODate("2026-09-20T08:41:00Z"),
  changedBy: "maria",
  changes: [
    { path: "title", from: "Launch notes", to: "Launch Notes" },
    { path: "tags", from: ["release"], to: ["release", "product"] }
  ]
}

This format is great for audit screens, because "what changed" is exactly what people want to see. Computing the diff in application code is straightforward for flat documents:

function diff(before, after, prefix = "") {
  const changes = [];
  const keys = new Set([
    ...Object.keys(before ?? {}),
    ...Object.keys(after ?? {}),
  ]);

  for (const key of keys) {
    const path = prefix ? `${prefix}.${key}` : key;
    const a = before?.[key];
    const b = after?.[key];

    const bothPlainObjects =
      a &&
      b &&
      typeof a === "object" &&
      typeof b === "object" &&
      !Array.isArray(a) &&
      !Array.isArray(b) &&
      !(a instanceof Date) &&
      !(b instanceof Date) &&
      a.constructor === Object &&
      b.constructor === Object;

    if (bothPlainObjects) {
      changes.push(...diff(a, b, path));
    } else if (JSON.stringify(a) !== JSON.stringify(b)) {
      changes.push({ path, from: a ?? null, to: b ?? null });
    }
  }
  return changes;
}

The trade-off is on the read side. To see what a document looked like at version 2, you start from the current version and apply diffs in reverse, or from an initial snapshot and apply diffs forward. A common hybrid stores a full snapshot every N versions (say, every 20) and diffs in between, so reconstruction never replays more than 19 changes.

Capturing History With Change Streams

The patterns above all put history-writing in your application code, which means every code path that writes the collection must remember to do it. Scripts, admin tools, and that one-off migration someone ran in mongosh all bypass it. Change streams move history capture out of the write path entirely.

A change stream is a real-time feed of changes to a collection, backed by the oplog. To get the previous version of a document in each event, enable pre- and post-images on the collection (MongoDB 6.0 and later):

db.runCommand({
  collMod: "articles",
  changeStreamPreAndPostImages: { enabled: true },
});

Then run a small worker that listens and writes history:

import { MongoClient } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
const db = client.db("cms");
const history = db.collection("articles_history");
const state = db.collection("stream_state");

async function run() {
  const saved = await state.findOne({ _id: "articles-history" });

  const stream = db.collection("articles").watch(
    [
      {
        $match: {
          operationType: { $in: ["insert", "update", "replace", "delete"] },
        },
      },
    ],
    {
      fullDocument: "whenAvailable",
      fullDocumentBeforeChange: "whenAvailable",
      ...(saved?.resumeToken && { resumeAfter: saved.resumeToken }),
    },
  );

  for await (const event of stream) {
    await history.insertOne({
      articleId: event.documentKey._id,
      op: event.operationType,
      before: event.fullDocumentBeforeChange ?? null,
      after: event.fullDocument ?? null,
      updatedFields: event.updateDescription?.updatedFields ?? null,
      removedFields: event.updateDescription?.removedFields ?? null,
      clusterTime: event.clusterTime,
      wallTime: event.wallTime,
    });

    await state.updateOne(
      { _id: "articles-history" },
      { $set: { resumeToken: event._id } },
      { upsert: true },
    );
  }
}

run().catch((err) => {
  console.error(err);
  process.exit(1);
});

The resume token is what makes this reliable. If the worker restarts, it picks up exactly where it left off, as long as the oplog still contains that point. Size your oplog (or Atlas oplog window) to comfortably cover your longest expected outage.

Change streams have limits you should be honest about. They capture what changed but not who or why, because that context lives in your application. A common workaround is to have the application set updatedBy and changeReason fields on every write, so they show up in the post-image. Also, history is written slightly after the change (eventually consistent), and if the worker is down longer than the oplog window, you lose events. For compliance-critical history, the transactional approach in Pattern 2 is stronger. For broad coverage of all writers, change streams win. Many teams run both. The post on building an event-driven audit log with change streams goes deeper on running these workers in production.

Querying History

With a history collection indexed on { articleId: 1, version: -1 }, the common queries are fast and simple.

List all versions of a document, newest first:

db.articles_history
  .find({ articleId: ObjectId("66f0c1a2e4b0a1b2c3d4e5f6") }, { snapshot: 0 })
  .sort({ version: -1 });

Get a specific version:

db.articles_history.findOne({
  articleId: ObjectId("66f0c1a2e4b0a1b2c3d4e5f6"),
  version: 2,
});

Find what a document looked like at a point in time. With superseded-version history, that's the earliest history entry whose supersededAt is after the target time; if none exists, the current version was already in effect:

db.articles_history
  .find({
    articleId: ObjectId("66f0c1a2e4b0a1b2c3d4e5f6"),
    changedAt: { $lte: ISODate("2026-09-01T00:00:00Z") },
    supersededAt: { $gt: ISODate("2026-09-01T00:00:00Z") },
  })
  .limit(1);

Everything a particular user changed this week, for an activity feed:

db.articles_history
  .find({ supersededBy: "sam", supersededAt: { $gte: ISODate("2026-09-21") } })
  .sort({ supersededAt: -1 });

Restoring a Previous Version

Restoring should be a new version, not a rewind. If an editor restores version 2 of an article that's currently on version 5, the result should be version 6 with version 2's content. That keeps history linear and means the restore itself is auditable:

export async function restoreVersion(id, targetVersion, user) {
  const old = await history.findOne({
    articleId: new ObjectId(id),
    version: targetVersion,
  });
  if (!old) throw new Error(`Version ${targetVersion} not found`);

  const current = await articles.findOne({ _id: new ObjectId(id) });
  return updateArticle(
    id,
    current.version,
    old.snapshot,
    user,
    `Restored version ${targetVersion}`,
  );
}

Reusing updateArticle means the restore gets the same transaction, concurrency check, and history entry as any other edit.

Retention and Storage

History grows forever unless you stop it. If you only need 180 days, a TTL index removes old entries automatically:

db.articles_history.createIndex(
  { supersededAt: 1 },
  { expireAfterSeconds: 60 * 60 * 24 * 180 },
);

See TTL indexes for expiring old data for the details and caveats. If you need long retention but rarely read old entries, consider moving history to a cheaper tier, such as Atlas Online Archive, or to a separate cluster.

For tamper resistance, give your application a database user that can only insert and find on the history collection, never update or remove. That won't stop a database administrator, but it prevents application bugs and compromised app credentials from rewriting the past.

Mongoose Users

If you use Mongoose, you can implement the history-collection pattern in middleware, but be careful: query middleware like pre("findOneAndUpdate") doesn't have the document, only the query, and updateMany touches many documents at once. The most reliable approach is an explicit service function like updateArticle above, or a change stream worker, rather than relying on hooks that some code paths skip. Mongoose's built-in __v field is for array concurrency, not general optimistic locking. If you want Mongoose to reject stale saves, enable optimisticConcurrency: true in the schema options, which applies to save().

Common Pitfalls

Writing history outside a transaction. Updating the document and then inserting history as two independent writes leaves a window where one succeeds and the other doesn't. Use a transaction, or accept eventual consistency with a change stream.

Unbounded embedded revisions. An array of revisions that grows forever eventually slows every read and can hit the 16 MB limit. Cap it with $slice or move history out.

Last-write-wins overwrites. Without a version check, two editors can each save over the other's work and your history will faithfully record the loss. Filter updates on the expected version.

Losing the "who." Neither the oplog nor change streams know which user made a change. Record updatedBy on the document itself, or in the history entry written by your application.

Restoring by overwriting history. Rewinding a document to an old version and deleting later revisions destroys the audit trail. Always restore forward as a new version.

Forgetting schema changes. A snapshot from two years ago may use fields your code no longer understands. Store a schemaVersion with each snapshot so restore logic can migrate old shapes.

Conclusion

MongoDB won't keep old versions for you, but a revision is just a document, so adding history is mostly a matter of choosing where it lives. Embedded revisions work for a short, capped trail. A separate history collection written in a transaction, with an optimistic version check, is the solid default for most applications. Diffs save space and power audit views, and change streams capture every write regardless of which code path made it.

Pick one collection in your app where "what did this say before?" has already come up, add a version field and an _history collection, and route its updates through a single transactional function. Once that works, you'll have a template you can reuse everywhere history matters.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading