
Point-in-Time Recovery in MongoDB Explained
It's 2:47 p.m. A script meant for staging runs against production and deletes every order older than 30 days. Your nightly backup ran at 2:00 a.m. Restoring it would bring the orders back, but it would also erase twelve hours of new orders, payments, and customer signups that happened since. You don't want last night. You want 2:46 p.m.
That's what point-in-time recovery (PITR) gives you: the ability to restore a database to any specific moment, not just the moments when backups happened to run. In MongoDB, it's built on two pieces you already have: a base backup and the oplog, the replication log that records every write in order.
This guide covers how point-in-time recovery works conceptually, how to use it in MongoDB Atlas with Continuous Cloud Backup, how to do it yourself on a self-managed replica set with mongodump and mongorestore, how to find the exact moment to restore to, and the limitations you need to plan around.
The Core Idea: Snapshot Plus Replay
Every point-in-time recovery system follows the same recipe:
- Restore a base backup taken before the target time. This gives you the database as it was at, say, 2:00 a.m.
- Replay the log of changes from the backup time forward, stopping just before the target time. In MongoDB, that log is the oplog.
The result is the database exactly as it was at 2:46:59 p.m., including every write that happened after the backup, but excluding the bad delete and everything after it.
The oplog works for this because every entry is idempotent and ordered by a timestamp. Each entry records the resulting change (for example, "set status to shipped on document X"), not the original command, so replaying it produces the same state no matter how many times it's applied. If you're curious about the details of those entries, MongoDB Oplog Explained: How Replication Works Under the Hood goes deep.
Why You Can't Just Use the Live Oplog
The oplog is a capped collection: it has a fixed size and overwrites its oldest entries as new ones arrive. On a busy cluster, it might only hold a few hours of history, known as the oplog window. For PITR, you need an unbroken chain of oplog entries from the base backup to the target time. If the oplog rolled over in between, the chain is broken and you can't replay across the gap.
That's why real PITR systems continuously copy the oplog somewhere durable, rather than relying on whatever happens to still be in local.oplog.rs when disaster strikes.
RPO: What PITR Changes
With snapshots alone, your recovery point is the last snapshot. With PITR, your recovery point can be seconds before the incident, limited only by how frequently the oplog is shipped to durable storage.
| Strategy | Can restore to | Worst-case data loss |
|---|---|---|
| Nightly snapshots | 2:00 a.m. each day | Up to 24 hours |
| Hourly snapshots | The top of each hour | Up to 1 hour |
| Snapshots plus PITR | Any second in the window | Seconds |
PITR in MongoDB Atlas
On Atlas dedicated clusters, PITR is a checkbox: Continuous Cloud Backup. When enabled, Atlas keeps taking scheduled snapshots and also streams the oplog to backup storage continuously. You configure a restore window (for example, the last 7 days), and you can restore to any moment within it.
Enabling It
In the Atlas UI, open your cluster's Backup settings and turn on Continuous Cloud Backup, then set the restore window in the backup policy. It adds cost, both for oplog storage and because a longer window requires keeping snapshots that cover it, so choose a window that matches how quickly you'd realistically notice a problem. For many teams, 3 to 7 days is a sensible default; data corruption bugs that go unnoticed for weeks need longer retention or separate long-term snapshots.
Running a Point-in-Time Restore
From the Backup tab, choose Restore, then Point in Time. You can specify the target either as a date and time, or as an oplog timestamp (seconds plus an increment), which lets you target a specific operation precisely.
With the Atlas CLI, a point-in-time restore to a separate cluster looks like this:
atlas backups restores start pointInTime \
--clusterName prod-cluster \
--pointInTimeUTCSeconds 1790606819 \
--targetClusterName prod-recovery \
--targetProjectId 64f1c0a2e4b0a1b2c3d4e5f6
Or, targeting an exact oplog position:
atlas backups restores start pointInTime \
--clusterName prod-cluster \
--oplogTs 1790606821 \
--oplogInc 3 \
--targetClusterName prod-recovery \
--targetProjectId 64f1c0a2e4b0a1b2c3d4e5f6
Check atlas backups restores start pointInTime --help for the exact flags on your CLI version. Then watch progress:
atlas backups restores list prod-cluster
Restore to a New Cluster First
Atlas can restore directly onto the source cluster, but doing that replaces all its data, including the good writes that happened after your target time. For a partial incident (one collection damaged, the rest fine), restore to a separate cluster and copy just the affected data back. That's covered in the section on surgical recovery below.
PITR on a Self-Managed Replica Set
Without Atlas, you can build the same capability with MongoDB Database Tools, or use a dedicated tool like Percona Backup for MongoDB (PBM) or Ops Manager. Building it by hand is worth understanding even if you end up using a tool, because every tool follows these steps.
Step 1: Take a Base Backup with the Oplog
Take a full dump that captures the oplog during the dump, so the base itself is consistent:
mongodump \
--uri="mongodb://backup:secret@db3.internal:27017/?authSource=admin&directConnection=true" \
--oplog \
--gzip \
--out=/backups/base-2026-09-28
Filesystem snapshots work as a base too, and are much faster for large datasets.
Step 2: Continuously Archive the Oplog
Between base backups, periodically dump new oplog entries. Each run picks up where the last one left off, using the timestamp of the last entry you saved:
#!/usr/bin/env bash
# archive-oplog.sh: run every few minutes from cron or a systemd timer
set -euo pipefail
URI="mongodb://backup:secret@db3.internal:27017/?authSource=admin&directConnection=true"
STATE=/backups/oplog/last_ts.json # e.g. {"t":1790582400,"i":1}
OUT=/backups/oplog/$(date -u +%Y%m%dT%H%M%SZ)
LAST=$(cat "$STATE")
mongodump --uri="$URI" \
--db=local --collection=oplog.rs \
--query="{\"ts\": {\"\$gt\": {\"\$timestamp\": $LAST}}}" \
--out="$OUT"
# Record the newest ts we just captured for the next run
mongosh "$URI" --quiet --eval '
const e = db.getSiblingDB("local").oplog.rs
.find({}, { ts: 1 }).sort({ $natural: -1 }).limit(1).next();
print(JSON.stringify({ t: e.ts.getHighBits(), i: e.ts.getLowBits() }));
' > "$STATE"
This is a simplified sketch: in production you'd upload each chunk to object storage, handle failures carefully, and verify there are no gaps between chunks. The key invariant is that the archived oplog must overlap the base backup and be continuous up to the present. The run interval sets your RPO, and it must be comfortably shorter than your oplog window.
Step 3: Find the Bad Operation
When something goes wrong, you need the timestamp of the first harmful operation. Search the oplog (live or archived) for it. For our accidental delete:
db.getSiblingDB("local")
.oplog.rs.find({ ns: "shop.orders", op: "d" }, { ts: 1, wall: 1, o: 1 })
.sort({ $natural: 1 })
.limit(3);
[
{
ts: Timestamp({ t: 1790606821, i: 3 }),
wall: ISODate("2026-09-28T14:47:01.244Z"),
o: { _id: ObjectId("66e1a0c4f1b2c3d4e5f60001") },
},
...
]
A deleteMany appears as many individual d (delete) entries, one per document, so look for the first one in the burst. A dropped collection appears as a command entry:
db.getSiblingDB("local").oplog.rs.find({
op: "c",
"o.drop": { $exists: true },
});
The wall field shows the wall-clock time, which helps you match the entry to what you saw in application logs. The ts value is what you'll actually use as the stopping point.
Step 4: Restore the Base Backup
Restore into a separate, isolated instance, never directly over production:
mongorestore \
--uri="mongodb://restore:secret@recovery.internal:27017/?authSource=admin" \
--gzip \
--oplogReplay \
--drop \
/backups/base-2026-09-28
Step 5: Replay the Oplog Up to the Target
mongorestore --oplogReplay expects a file named oplog.bson at the top of the directory it's given. Combine your archived chunks from after the base backup into one ordered file, then replay with --oplogLimit, which applies entries up to but not including the given timestamp:
mkdir -p /restore/replay
# Concatenate archived oplog chunks in chronological order
cat /backups/oplog/2026*/local/oplog.rs.bson > /restore/replay/oplog.bson
mongorestore \
--uri="mongodb://restore:secret@recovery.internal:27017/?authSource=admin" \
--oplogReplay \
--oplogLimit=1790606821:3 \
/restore/replay
2026-09-28T15:20:41.901+0000 replaying oplog
2026-09-28T15:21:06.337+0000 applied 1843220 oplog entries
2026-09-28T15:21:06.338+0000 done
Because oplog entries are idempotent, overlap between the base backup's captured oplog and the first archived chunk is harmless. Gaps are not: if a chunk is missing, the replay produces an inconsistent result.
The recovery instance now reflects production as of the moment just before Timestamp(1790606821, 3).
Surgical Recovery: Copying Back Only What Was Lost
A full rollback throws away every legitimate write since the incident. Usually, you only need to recover the data that was damaged. With a recovered copy running separately, you can compare and copy back:
// Connected to production with access to the recovery instance's data,
// or export from recovery and import into a staging collection first.
const recovered = db.getSiblingDB("recovery_shop").orders;
const live = db.getSiblingDB("shop").orders;
const cutoff = new Date(Date.now() - 30 * 24 * 60 * 60 * 1000);
let restored = 0;
recovered.find({ createdAt: { $lt: cutoff } }).forEach((doc) => {
const res = live.replaceOne({ _id: doc._id }, doc, { upsert: true });
if (res.upsertedCount) restored++;
});
print(`restored ${restored} orders`);
In practice, you'd mongodump the affected collection from the recovery instance, mongorestore it into a staging namespace on production (using --nsFrom and --nsTo), then merge with a script like the one above. For large volumes, batch the writes with bulkWrite().
Be careful with documents that were legitimately changed after the incident. If an order was deleted accidentally but a related refund was recorded afterward, a naive copy-back might miss the relationship. Think through the business rules before merging.
Sharded Clusters
PITR on sharded clusters is significantly harder. Each shard has its own oplog, and restoring to a consistent point requires coordinating all shards and the config servers to the same cluster time. Chunk migrations during the window add further complexity.
Doing this by hand isn't practical. Use Atlas Continuous Cloud Backup, Ops Manager, or Percona Backup for MongoDB, all of which support consistent point-in-time restores of sharded clusters. With PBM, for example, PITR is enabled with a configuration flag and a restore targets a timestamp:
pbm config --set pitr.enabled=true
pbm restore --time="2026-09-28T14:46:59"
Limitations and Gotchas
The oplog must be continuous. Any gap between the base backup and the target time makes PITR to that point impossible. Monitor your archiving job and alert if it falls behind.
Your window is bounded by retention. You can't restore to a point before your oldest retained base backup, or beyond your oplog retention.
Some operations are coarse in the oplog. Collection drops and index builds are single command entries. That's fine for replay, but it means you stop before the drop, not "halfway through" it.
Restores take time. Replaying days of oplog on a busy cluster can take hours. Test to learn your real restore time, and take base backups often enough to keep replay short.
Time zones cause mistakes. Oplog timestamps are in seconds since the Unix epoch (UTC), and wall values are UTC. Convert application log times carefully before choosing a target.
Recovery is a data problem, not just a restore. Getting a correct copy is only half the work. Merging recovered data back into a live system needs thought about what changed afterward.
Conclusion
Point-in-time recovery combines a base backup with a continuous, replayable log of changes, letting you restore MongoDB to the second before an incident rather than to whenever the last backup happened to run. Atlas makes it a setting with Continuous Cloud Backup. On self-managed replica sets, you can build it with mongodump --oplog, continuous oplog archiving, and mongorestore --oplogReplay --oplogLimit, or adopt a tool like Percona Backup for MongoDB, especially for sharded clusters.
Your next step: run a PITR drill. Insert a known document into a test cluster, delete it, and then recover to the moment just before the delete using either the Atlas point-in-time restore or the manual steps above. Once you've done it on a calm day, you'll be ready when you need it at 2:47 p.m.


