Type something to search...
Backup and Restore Strategies for MongoDB: mongodump, Snapshots, and Atlas Backups

Backup and Restore Strategies for MongoDB: mongodump, Snapshots, and Atlas Backups

Replica sets protect you from a server dying. They don't protect you from a developer running deleteMany({}) against the wrong database, a migration script with a bug, ransomware, or an application that slowly corrupts records for a week before anyone notices. Replication copies mistakes to every member within milliseconds. Only backups let you go back in time.

MongoDB gives you several ways to take backups, and they differ enormously in speed, consistency, and how much operational effort they need. A logical dump with mongodump is simple and portable but slow on large datasets. Filesystem snapshots are fast but need the right storage setup. MongoDB Atlas handles backups for you, with scheduling, retention, and point-in-time restores built in. The right choice depends on your data size, your deployment, and how much data you can afford to lose.

This guide covers the three main approaches, how to run each one correctly, how to restore from them, and how to put together a backup strategy you've actually verified.

Start with RPO and RTO

Before picking a tool, answer two questions:

  • Recovery Point Objective (RPO): How much data can you afford to lose? If you back up nightly, a failure at 11 p.m. loses almost a day of writes.
  • Recovery Time Objective (RTO): How long can you be down while restoring? Restoring 2 TB from a logical dump can take many hours; restoring from a snapshot can take minutes.

These numbers drive everything else. An internal analytics database might tolerate a 24-hour RPO and a few hours of RTO. A payment system might need an RPO of seconds and an RTO under an hour, which rules out nightly dumps entirely and points toward continuous backups with point-in-time recovery.

Option 1: mongodump and mongorestore

mongodump and mongorestore are part of the MongoDB Database Tools, installed separately from the server. mongodump connects like a regular client, reads documents, and writes them out as BSON files (plus metadata for indexes and collection options). mongorestore reads those files back and inserts them.

Basic Dump

mongodump \
  --uri="mongodb://backup:secret@db1.internal:27017/?authSource=admin&replicaSet=rs0&readPreference=secondary" \
  --out=/backups/2026-09-23
2026-09-23T02:00:04.118+0000    writing shop.orders to /backups/2026-09-23/shop/orders.bson
2026-09-23T02:00:04.120+0000    writing shop.customers to /backups/2026-09-23/shop/customers.bson
2026-09-23T02:03:51.402+0000    done dumping shop.orders (4812330 documents)

Reading from a secondary keeps the dump from competing with production traffic on the primary.

For backups, a single compressed archive is usually easier to handle than a directory tree:

mongodump \
  --uri="mongodb://backup:secret@db1.internal:27017/?authSource=admin&replicaSet=rs0&readPreference=secondary" \
  --archive=/backups/shop-2026-09-23.archive.gz \
  --gzip \
  --oplog

Why --oplog Matters

A dump of a busy database takes time, and writes continue while it runs. Without extra care, the collections you dump first reflect an earlier moment than the ones you dump last. The --oplog flag fixes this for replica sets: mongodump also captures the oplog entries written during the dump. On restore, mongorestore --oplogReplay applies those entries, bringing every collection to the same point in time: the moment the dump finished.

--oplog works only for full-instance dumps of a replica set member (not with --db or --collection). It isn't supported for sharded clusters through mongos.

The backup user needs the built-in backup role:

db.getSiblingDB("admin").createUser({
  user: "backup",
  pwd: passwordPrompt(),
  roles: [{ role: "backup", db: "admin" }],
});

Restoring

Restore the archive into a cluster, replaying the captured oplog:

mongorestore \
  --uri="mongodb://restore:secret@staging1.internal:27017/?authSource=admin" \
  --archive=/backups/shop-2026-09-23.archive.gz \
  --gzip \
  --oplogReplay \
  --drop

--drop drops each collection before restoring it, so you end up with exactly the backed-up data instead of a merge. The restoring user needs the restore role.

You can also restore selectively. To pull back just one collection into a different namespace (handy when you need to recover a few documents without overwriting the live collection):

mongorestore \
  --uri="mongodb://restore:secret@db1.internal:27017/?authSource=admin" \
  --archive=/backups/shop-2026-09-23.archive.gz \
  --gzip \
  --nsInclude="shop.orders" \
  --nsFrom="shop.orders" \
  --nsTo="recovery.orders_20260923"

Then copy the needed documents from recovery.orders_20260923 back into shop.orders with a targeted query.

When mongodump Fits

Strengths: simple, works with any deployment you can connect to, portable across versions and platforms, easy to restore selectively.

Weaknesses: slow for large datasets (it reads every document through the query layer), puts load on the server, rebuilds all indexes on restore (which can take longer than the data load), and isn't a consistent backup method for sharded clusters.

A good rule of thumb: mongodump is excellent for small to medium databases (tens of gigabytes), development and staging snapshots, and migrating individual databases. For large production datasets, it's a supplement, not the primary strategy.

Option 2: Filesystem Snapshots

A filesystem or block-level snapshot captures the data files as they exist at an instant. On cloud providers, this means an EBS snapshot, a persistent disk snapshot, or an Azure managed disk snapshot. On-premises, LVM or storage-array snapshots do the same job.

Snapshots are fast to take (seconds, regardless of data size) and fast to restore, because you attach a volume rather than reinserting documents and rebuilding indexes.

Consistency Rules

WiredTiger is crash-consistent: if a snapshot captures the data files and the journal at the same instant, mongod recovers from it just like it would after a power loss. That leads to one crucial requirement: the data files and the journal must be on the same volume, captured by the same snapshot.

If they're on separate volumes, or you're snapshotting multiple volumes that can't be captured atomically, you need to pause writes during the snapshot:

// On the member being snapshotted (ideally a hidden or dedicated secondary)
db.fsyncLock();
aws ec2 create-snapshot \
  --volume-id vol-0a1b2c3d4e5f67890 \
  --description "mongo rs0 member3 2026-09-23"
db.fsyncUnlock();

db.fsyncLock() flushes pending writes and blocks new ones on that member. Always run it on a secondary, never the primary, and make sure your script unlocks even if the snapshot command fails.

A Dedicated Backup Member

A common pattern is to add a hidden secondary with priority 0 that exists purely for backups:

rs.add({
  host: "db-backup.internal:27017",
  priority: 0,
  hidden: true,
  votes: 0,
});

It replicates everything but never becomes primary and never serves application reads, so snapshotting it (or locking it briefly) has no user-facing impact.

Restoring from a Snapshot

To restore, create a new volume from the snapshot, attach it to a server, and start mongod with that volume as dbPath. For a full replica set restore, you'd typically bring up one member from the snapshot, then either restore the same snapshot to the other members or let them initial-sync from it. MongoDB's documentation has detailed procedures for restoring replica sets and sharded clusters from snapshots; follow them closely, because steps like resetting replica set configuration matter.

Sharded Clusters

Snapshots of a sharded cluster need to be coordinated: you must stop the balancer and capture every shard and the config server replica set at a consistent moment. Doing this by hand is fragile. For sharded clusters, a purpose-built tool (Atlas backups, Ops Manager or Cloud Manager, or Percona Backup for MongoDB) is strongly recommended.

Option 3: Atlas Cloud Backups

If you're on MongoDB Atlas with a dedicated cluster (M10 and up), backups are a managed feature. Atlas takes cloud backup snapshots using your cloud provider's native snapshot mechanism, on a schedule you define, with separate retention for hourly, daily, weekly, monthly, and yearly snapshots.

The free M0 tier doesn't include cloud backups, and Flex clusters include a more limited daily snapshot. For production, a dedicated tier with a configured backup policy is the norm.

Continuous Cloud Backup

Enabling Continuous Cloud Backup adds point-in-time recovery: Atlas captures the oplog continuously alongside snapshots, so you can restore to any moment within the configured restore window, down to the second. That shrinks your RPO from "since the last snapshot" to effectively zero. Point-in-Time Recovery in MongoDB Explained covers how this works in detail.

Managing Backups with the Atlas CLI

List snapshots for a cluster:

atlas backups snapshots list prod-cluster --output json

Take an on-demand snapshot before a risky migration:

atlas backups snapshots create prod-cluster \
  --desc "Before orders schema migration 2026-09-23"

Restore a snapshot to a different cluster (the safest way to recover, because it doesn't overwrite production):

atlas backups restores start automated \
  --clusterName prod-cluster \
  --snapshotId 66f0b8e2a1c3d4e5f6a7b8c9 \
  --targetClusterName restore-check \
  --targetProjectId 64f1c0a2e4b0a1b2c3d4e5f6

You can also download a snapshot for local use or archival, or restore directly onto the source cluster when you're truly rolling back everything.

Compliance and Protection

Atlas offers a Backup Compliance Policy at the project level, which prevents anyone (including project owners) from deleting backups or shortening retention below a defined minimum. That's a meaningful defense against ransomware and malicious insiders, since an attacker with admin access can't simply delete your safety net. Atlas also supports copying snapshots to other regions for disaster recovery.

Other Tools for Self-Managed Clusters

For self-managed deployments that outgrow mongodump and hand-rolled snapshots:

  • Ops Manager / Cloud Manager (MongoDB, requires an Enterprise Advanced subscription for Ops Manager) provide managed backups with point-in-time restore for replica sets and sharded clusters.
  • Percona Backup for MongoDB (PBM) is open source and supports consistent backups of replica sets and sharded clusters, including physical backups and point-in-time recovery, stored in S3-compatible object storage.

Comparing the Options

ApproachSpeed on large dataConsistencySharded supportPoint-in-timeEffort
mongodumpSlowWith --oplog (RS)Not consistentNoLow
Filesystem snapshotsFastCrash-consistentNeeds coordinationWith extra toolingMedium
Atlas cloud backupsFastManagedYesYes (continuous)Very low
PBM / Ops ManagerFast (physical)ManagedYesYesMedium

Many teams combine approaches: Atlas continuous backups or snapshots as the primary mechanism, plus periodic mongodump exports of critical databases stored in a separate account or provider as a portable, independent copy.

Best Practices

Test restores on a schedule. A backup is only proven when you've restored it. Automate a monthly restore into a scratch cluster and run a sanity check: document counts, a few known records, index presence.

// Quick sanity check after a restore
const dbs = ["shop", "billing"];
dbs.forEach((name) => {
  const d = db.getSiblingDB(name);
  d.getCollectionNames().forEach((c) => {
    print(`${name}.${c}: ${d.getCollection(c).estimatedDocumentCount()}`);
  });
});

Follow the 3-2-1 rule. Keep at least three copies of your data, on two different types of storage, with one copy off-site (a different region, account, or provider). A backup stored in the same cloud account as production can be deleted by the same compromised credentials.

Back up from secondaries. Dumps and snapshots add load. A hidden secondary dedicated to backups keeps that load away from users.

Encrypt and restrict access. Backups contain all your data. Encrypt archives at rest, restrict who can read the bucket, and use dedicated users with the backup and restore roles rather than root credentials.

Include users, roles, and configuration. A full-instance mongodump includes the admin database with users and roles. Selective dumps don't. Keep your mongod.conf, replica set configuration, and index definitions in version control too.

Monitor backup jobs. A cron job that silently stopped three months ago is a common disaster story. Alert on backup failures and on the age of the newest successful backup.

Take a manual snapshot before risky changes. Migrations, major version upgrades, and bulk data fixes all deserve an on-demand backup first. It costs little and turns a potential disaster into a short rollback.

Conclusion

MongoDB backups come in three main flavors. mongodump with --oplog is simple and portable, ideal for smaller databases and selective recovery. Filesystem snapshots are fast and scale to large datasets, as long as the journal and data share a volume or you lock writes on a secondary. Atlas cloud backups remove the operational burden entirely and add point-in-time recovery with Continuous Cloud Backup. For self-managed sharded clusters, reach for a dedicated tool like Percona Backup for MongoDB or Ops Manager.

Whatever you choose, the strategy isn't complete until you've restored from it. Your next step: pick yesterday's backup, restore it into a throwaway cluster, time how long it takes, and compare that number to your RTO. If it's longer than you expected, you've learned something important on a quiet day instead of during an outage.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading