Type something to search...
Database Migrations in MongoDB: Tools and Strategies for Evolving Data

Database Migrations in MongoDB: Tools and Strategies for Evolving Data

One of MongoDB's selling points is that you don't have to run ALTER TABLE to add a field. That's true, and it's genuinely useful early in a project. But six months later, your users collection has documents where the name is a single string, documents where it's split into firstName and lastName, and a handful where it's missing entirely because of a bug in a signup form that shipped for two days. Every query and every piece of code that reads users now has to cope with all three shapes.

The schema didn't go away; it just moved into your application code, where it's harder to see. Migrations are how you keep that implicit schema under control: versioned, repeatable scripts that move existing data from one shape to the next, plus the index and validation changes that go with it. MongoDB gives you more flexibility than a relational database in when you migrate, but you still need a plan.

This guide covers the two main strategies (eager and lazy migration), the schema versioning pattern, using migrate-mongo to run versioned scripts, writing migrations that are safe on large collections, handling index and validator changes, and deploying schema changes without downtime.

Why You Still Need Migrations

In a relational database, the schema is enforced, so a migration is mandatory before code can use a new column. In MongoDB, the database will happily accept documents of any shape, which creates a different set of problems:

  • Queries silently miss documents. A query on { "address.city": "Lisbon" } doesn't match users who still have a flat city field. No error, just wrong results.
  • Code accumulates branches. Every reader grows fallbacks like user.fullName ?? user.firstName + " " + user.lastName, and each new reader has to rediscover them.
  • Indexes and validation drift. A unique index added in production but not in staging, or a $jsonSchema validator that exists in one environment only.
  • Reporting breaks. Aggregations that group by a field that's been renamed quietly lump old documents into a null bucket.

A migration process fixes all of these by making changes explicit, versioned, and repeatable in every environment.

Two Strategies: Eager and Lazy

Eager Migration

An eager migration rewrites all affected documents at once, usually as part of a deploy. After it runs, every document has the new shape and the application only needs to understand one version.

// Split "name" into firstName / lastName for every user that still has it
db.users.updateMany(
  { name: { $exists: true }, firstName: { $exists: false } },
  [
    {
      $set: {
        firstName: { $arrayElemAt: [{ $split: ["$name", " "] }, 0] },
        lastName: {
          $trim: {
            input: {
              $reduce: {
                input: { $slice: [{ $split: ["$name", " "] }, 1, 100] },
                initialValue: "",
                in: { $concat: ["$$value", " ", "$$this"] },
              },
            },
          },
        },
      },
    },
    { $unset: "name" },
  ],
);

This uses an update with an aggregation pipeline (the array as the second argument), which lets you compute new fields from existing ones entirely on the server. No documents travel to your app and back.

Eager migrations are simple to reason about, but on a large collection they generate a lot of write load and oplog traffic, and they can take a long time.

Lazy Migration

A lazy migration (sometimes called migrate-on-read) leaves existing documents alone and upgrades each one the next time the application reads or writes it. New documents are written in the new shape from day one.

function upgradeUser(doc) {
  if (doc.schemaVersion === 2) return doc;
  const [firstName, ...rest] = (doc.name ?? "").split(" ");
  const { name, ...others } = doc;
  return { ...others, firstName, lastName: rest.join(" "), schemaVersion: 2 };
}

async function getUser(id) {
  const raw = await users.findOne({ _id: id });
  if (!raw) return null;
  const user = upgradeUser(raw);
  if (raw.schemaVersion !== 2) {
    await users.replaceOne({ _id: id, schemaVersion: raw.schemaVersion }, user);
  }
  return user;
}

The schemaVersion in the replaceOne filter acts as an optimistic concurrency check: if another request upgraded the document first, this replace simply matches nothing.

Lazy migration spreads the cost over time and avoids a big-bang write. The downside is that your code has to understand every version still in the database, and documents that are rarely read may stay on the old shape forever. Queries that filter on the new field won't find unmigrated documents.

Which One to Use

FactorEagerLazy
Collection sizeSmall to mediumVery large
Queries depend on new shapeRequiredRisky until complete
Code complexityLow after migrationHigher, multiple versions
Write loadConcentratedSpread out
ReversibilityNeeds a down scriptOld docs untouched until read

In practice, many teams combine them: deploy code that reads both shapes and writes the new one (lazy), then run a background eager migration to convert the stragglers, then remove the old-shape code.

The Schema Versioning Pattern

Whatever strategy you choose, adding a schemaVersion field to documents makes migrations much easier to manage. It tells both your code and your migration scripts exactly what shape a document is in, instead of inferring it from which fields exist.

db.users.insertOne({
  schemaVersion: 2,
  firstName: "Ada",
  lastName: "Lovelace",
  email: "ada@example.com",
});

You can then find and count documents per version:

db.users.aggregate([{ $group: { _id: "$schemaVersion", count: { $sum: 1 } } }]);
[
  { "_id": null, "count": 18204 },
  { "_id": 2, "count": 311592 }
]

The null bucket shows documents created before versioning was introduced. When that count reaches zero, you can safely delete the code path that handles them.

Versioned Migration Scripts with migrate-mongo

For eager migrations, you want the same tooling relational databases have had for years: numbered scripts, a record of which ones have run, and a command to apply pending ones. migrate-mongo is the most widely used option in the Node.js ecosystem.

npm install --save-dev migrate-mongo
npx migrate-mongo init

This creates migrate-mongo-config.js and a migrations/ folder. Point the config at your database through an environment variable:

// migrate-mongo-config.js
module.exports = {
  mongodb: {
    url: process.env.MONGODB_URI,
    databaseName: process.env.MONGODB_DB,
  },
  migrationsDir: "migrations",
  changelogCollectionName: "changelog",
  lockCollectionName: "changelog_lock",
  lockTtl: 0,
  migrationFileExtension: ".js",
  useFileHash: false,
  moduleSystem: "commonjs",
};

Create a migration:

npx migrate-mongo create split-user-names

That generates a timestamped file such as migrations/20260914063400-split-user-names.js. Fill in up and down:

module.exports = {
  async up(db) {
    await db
      .collection("users")
      .updateMany({ schemaVersion: { $exists: false } }, [
        {
          $set: {
            firstName: { $arrayElemAt: [{ $split: ["$name", " "] }, 0] },
            lastName: {
              $reduce: {
                input: { $slice: [{ $split: ["$name", " "] }, 1, 100] },
                initialValue: "",
                in: {
                  $trim: { input: { $concat: ["$$value", " ", "$$this"] } },
                },
              },
            },
            schemaVersion: 2,
          },
        },
        { $unset: "name" },
      ]);
  },

  async down(db) {
    await db.collection("users").updateMany({ schemaVersion: 2 }, [
      {
        $set: {
          name: {
            $trim: { input: { $concat: ["$firstName", " ", "$lastName"] } },
          },
        },
      },
      { $unset: ["firstName", "lastName", "schemaVersion"] },
    ]);
  },
};

Then run it and check the status:

npx migrate-mongo up
npx migrate-mongo status
┌──────────────────────────────────────────┬──────────────────────────┐
│ Filename                                 │ Applied At               │
├──────────────────────────────────────────┼──────────────────────────┤
│ 20260914063400-split-user-names.js       │ 2026-09-14T06:41:12.338Z │
└──────────────────────────────────────────┴──────────────────────────┘

migrate-mongo records each applied migration in the changelog collection, so running up again is a no-op. down reverts the most recent migration.

Other Options

  • Mongoose projects often use ts-migrate-mongoose, which runs migrations with your Mongoose models available. Be careful: if a model changes later, old migrations that import it may break. Using the raw db handle in migrations avoids this.
  • Java and Spring Boot teams commonly use Mongock, which integrates with Spring Data MongoDB.
  • Python teams frequently roll their own small runner with PyMongo, or use a library like pymongo-migrate. The core idea (a changelog collection plus ordered scripts) is simple enough to write yourself.
  • Liquibase also supports MongoDB through an extension if your organization standardizes on it.

Writing Migrations That Are Safe on Large Collections

A single updateMany over 200 million documents will work, eventually, but it can saturate disk I/O, flood the oplog, increase replication lag, and hold a long-running operation that's hard to monitor. For big collections, migrate in batches.

// migrations/20260920090000-backfill-order-totals.js
module.exports = {
  async up(db) {
    const orders = db.collection("orders");
    const batchSize = 1000;
    let lastId = null;

    while (true) {
      const filter = { totalCents: { $exists: false } };
      if (lastId) filter._id = { $gt: lastId };

      const batch = await orders
        .find(filter, { projection: { _id: 1, total: 1 } })
        .sort({ _id: 1 })
        .limit(batchSize)
        .toArray();

      if (batch.length === 0) break;

      await orders.bulkWrite(
        batch.map((o) => ({
          updateOne: {
            filter: { _id: o._id, totalCents: { $exists: false } },
            update: { $set: { totalCents: Math.round(o.total * 100) } },
          },
        })),
        { ordered: false },
      );

      lastId = batch[batch.length - 1]._id;
      await new Promise((r) => setTimeout(r, 50)); // give replication room to breathe
    }
  },

  async down(db) {
    await db
      .collection("orders")
      .updateMany({}, { $unset: { totalCents: "" } });
  },
};

A few things make this safe:

  • Range-based iteration on _id instead of skip(), so each batch is an efficient index scan. The same idea as range-based pagination.
  • Idempotent filters. Each update includes totalCents: { $exists: false }, so rerunning the migration after a crash skips work already done.
  • A small pause between batches to keep replication lag in check. Watch rs.printSecondaryReplicationInfo() during big migrations and tune the batch size.
  • Unordered bulk writes so one bad document doesn't stop the batch.

For really heavy transformations, you can also write the reshaped data into a new collection with $merge or $out and then swap names, but that requires stopping writes during the switch.

Index and Validator Changes

Migrations aren't only about document shape. Index and validation changes need the same version control.

Creating Indexes

Since MongoDB 4.2, index builds use an optimized process that holds an exclusive lock only briefly at the start and end, so creating an index on a live collection is generally safe. It still uses significant resources on large collections, so run heavy builds at low-traffic times:

module.exports = {
  async up(db) {
    await db
      .collection("orders")
      .createIndex(
        { customerId: 1, createdAt: -1 },
        { name: "customer_recent_orders" },
      );
  },
  async down(db) {
    await db.collection("orders").dropIndex("customer_recent_orders");
  },
};

Adding a unique index will fail if duplicates already exist. Check first, and include deduplication in the migration if needed:

db.users.aggregate([
  {
    $group: {
      _id: { $toLower: "$email" },
      n: { $sum: 1 },
      ids: { $push: "$_id" },
    },
  },
  { $match: { n: { $gt: 1 } } },
]);

Replacing an Index Without a Gap

To change an index definition, create the new index first, deploy code that uses it, then drop the old one. Dropping first leaves a window where queries fall back to collection scans. You can also hide the old index with db.collection.hideIndex() to see how queries behave without it before dropping it permanently.

Updating Validators

If you use schema validation, change validators with collMod, and consider validationLevel: "moderate" during the transition so existing non-conforming documents can still be updated:

await db.command({
  collMod: "users",
  validator: {
    $jsonSchema: {
      bsonType: "object",
      required: ["email", "firstName", "schemaVersion"],
      properties: {
        email: { bsonType: "string" },
        firstName: { bsonType: "string" },
        schemaVersion: { bsonType: "int", minimum: 2 },
      },
    },
  },
  validationLevel: "moderate",
  validationAction: "error",
});

Once the data migration is complete, switch to validationLevel: "strict" in a follow-up migration.

Zero-Downtime Schema Changes: Expand and Contract

Renaming or restructuring a field that live code depends on can't happen in a single deploy without breaking something: either old code sees new data or new code sees old data. The expand and contract pattern (also called parallel change) solves this in stages:

  1. Expand. Deploy code that writes both the old and new fields, and reads the new field with a fallback to the old one.
  2. Migrate. Run a batched backfill so every document has the new field.
  3. Switch. Deploy code that reads only the new field and stops writing the old one.
  4. Contract. Run a migration that removes the old field, and drop any indexes that referenced it.

Each step is independently deployable and reversible, and at no point is running code confronted with data it doesn't understand. It takes longer than a single migration, but for anything customer-facing, it's worth it.

Running Migrations in Your Pipeline

Decide deliberately where migrations run:

  • As a separate deploy step (a CI job or a Kubernetes Job) before the new application version rolls out. This is the most common and the easiest to monitor.
  • On application startup. Convenient for small apps, but dangerous with multiple instances, which will race to run the same migration. migrate-mongo's lock collection helps, but long migrations will also delay startup and health checks.

Always take a backup (or confirm that point-in-time recovery is enabled) before running a destructive migration, and test the migration against a copy of production data. A migration that takes three seconds on a 10,000-document staging database may take three hours in production.

Common Mistakes

Assuming "schemaless" means no migrations. The schema lives in your code whether you manage it or not. Unmanaged, it becomes a pile of fallbacks and silently wrong queries.

Running one giant updateMany on a huge collection. Batch it, make it idempotent, and watch replication lag.

Importing application models into migrations. When the model changes next quarter, your old migration breaks. Use the raw driver in migration scripts.

Dropping the old field in the same deploy that introduces the new one. Old instances still running during a rolling deploy will break. Use expand and contract.

Skipping the down script. Even if you rarely roll back, writing down forces you to think about whether the change is reversible, and whether you're destroying information you can't recover.

Conclusion

MongoDB's flexible schema lets you choose how and when to evolve your data, but it doesn't remove the need to do it deliberately. Use eager migrations with batched, idempotent scripts when queries depend on the new shape, lazy migration with a schemaVersion field when collections are huge, and expand and contract whenever live code is involved. Keep it all in versioned scripts with a tool like migrate-mongo so every environment ends up in the same state.

As a first step, run a $group on schemaVersion (or on the presence of a field you know has changed) across your largest collection. If you find more than one shape, write your first migration to consolidate them.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading