
MongoDB Retryable Writes and Reads: Building Resilient Applications
In a healthy production cluster, things still go wrong all the time. A primary steps down for a rolling upgrade. A load balancer drops an idle connection. A cloud provider migrates a VM and the network pauses for two seconds. None of these are outages, but each one can make a database operation fail. If your application turns every one of those failures into a 500 error, users will notice, even though the database recovered almost immediately.
Retryable writes and retryable reads are MongoDB's built-in defense. When an operation fails with a transient network or failover error, the driver waits for a suitable server and automatically tries the operation one more time. For writes, the server guarantees that a retried operation is applied at most once, even if the first attempt actually succeeded before the connection dropped. Both features are on by default in every current driver.
This guide covers how retryable writes avoid duplicates, exactly which operations are retried and which aren't, how retryable reads work, how to test failure handling with fail points, and how to build application-level retries for everything the driver doesn't cover.
The Problem: Did My Write Happen?
Consider an insert that's sent to the primary. The primary applies it, but the network connection drops before the acknowledgment reaches your application. From the application's point of view, the operation failed with a network error. But the document exists.
If you retry naively, you might insert it twice. If you don't retry, you might report failure for something that succeeded. Without help from the database, there's no way to know which situation you're in.
For an $inc on an account balance, a duplicate is a real bug. For an order insert with a generated _id, you'll get two orders. This is the problem retryable writes solve.
How Retryable Writes Work
When retryable writes are enabled, the driver attaches two things to each eligible write:
- A logical session ID (
lsid), which identifies the client session. - A transaction number (
txnNumber), which increases with each write in that session.
The server records the outcome of each (lsid, txnNumber) pair, and that record is replicated along with the write itself. If the driver retries a write with the same pair, the server recognizes it:
- If the first attempt was applied, the server returns the original result without applying it again.
- If the first attempt wasn't applied, the server applies it now.
Because the record is replicated, this works across failovers. If the primary applies your write, replicates it, and then crashes before replying, the new primary knows about the (lsid, txnNumber) and returns the saved result when the driver retries.
The retry flow in the driver looks like this:
- Send the write. It fails with a retryable error (a network error, or a server error labeled
RetryableWriteError, such as "not primary" or "shutdown in progress"). - Run server selection again, waiting for a new primary if necessary (bounded by
serverSelectionTimeoutMS). - Send the same write, with the same
lsidandtxnNumber, once. - If the retry also fails, return the error to the application.
The driver retries exactly once. The goal is to cover a single transient event, like one election, not to keep hammering a broken cluster.
Enabling and Configuring
Retryable writes and reads are enabled by default in all current drivers (Node.js 6.x, PyMongo 4.x, and the rest). You'd have to explicitly disable them:
// These are the defaults; you don't need to set them
const client = new MongoClient(
"mongodb+srv://app:secret@cluster0.example.mongodb.net/?retryWrites=true&retryReads=true",
);
from pymongo import MongoClient
client = MongoClient(
"mongodb://db-1,db-2,db-3/?replicaSet=rs0",
retryWrites=True, # default
retryReads=True, # default
)
There are a few requirements:
- A replica set or sharded cluster. Standalone servers don't support retryable writes, since there's no oplog to record them in. Drivers silently skip retries on standalones. For local development, run a single-node replica set.
- Acknowledged writes. Writes with
w: 0can't be retried, because the driver never learns the outcome. - Not inside a transaction. Operations within a multi-document transaction aren't retried individually. Transactions have their own retry mechanism, covered below.
If you find retryWrites=false in an existing connection string, it's almost always left over from an old setup, and you should remove it.
Which Writes Are Retryable
The rule is: a write is retryable if it affects at most one document (or is an insert). Operations that can modify many documents aren't retryable, because the server can't cheaply record "which of these 50,000 documents did I already update?"
| Operation | Retryable? |
|---|---|
insertOne, insertMany | Yes |
updateOne, replaceOne | Yes |
deleteOne | Yes |
findOneAndUpdate, findOneAndReplace, findOneAndDelete | Yes |
bulkWrite with only single-document operations | Yes |
updateMany, deleteMany | No |
bulkWrite containing updateMany or deleteMany | No |
Writes with w: 0 | No |
For insertMany and bulkWrite, the driver splits the batch into commands and retries a failed command as a whole; already-applied documents in that command are recognized and not duplicated.
The gap to watch is multi-document updates and deletes. If an updateMany fails with a network error, the driver returns the error, and the update may have been partially applied. Some documents changed, some didn't.
Making Multi-Document Writes Safe to Retry
The fix is to make the operation idempotent, so running it again produces the same result. That usually means the filter excludes documents that were already updated:
// Not idempotent: running twice applies the discount twice
await products.updateMany({ category: "lamps" }, { $mul: { price: 0.9 } });
// Idempotent: a second run matches nothing already discounted
await products.updateMany(
{ category: "lamps", "promo.fall2026": { $ne: true } },
{ $mul: { price: 0.9 }, $set: { "promo.fall2026": true } },
);
Now if the first call fails partway through, you can safely run it again. Documents that were already discounted are skipped. The same idea applies to deletes: deleteMany({ expiresAt: { $lt: cutoff } }) is naturally idempotent, since deleting a missing document is a no-op.
Retryable Reads
Reads have a simpler problem: retrying a read never causes duplicates. Retryable reads just save you from surfacing a transient error. When a read fails with a network or "not primary" type error, the driver selects a server again and retries once.
Operations covered include find, findOne, aggregate (without $out or $merge), distinct, countDocuments, estimatedDocumentCount, watch (opening a change stream), and listing commands such as listCollections, listIndexes, and listDatabases.
Two details are worth knowing:
- Only the initial command is retried. For a
findthat returns many batches, the driver retries the first request, but if the connection drops while fetching a later batch withgetMore, the error is returned. The cursor is tied to a specific server and can't simply be resumed elsewhere. For long scans, design your code to restart from the last processed_id. - Aggregations that write aren't retried, because
$outand$mergemodify data.
Change streams go further than retryable reads: the driver automatically resumes them after transient errors using the last resume token, so a long-running stream survives failovers without any code on your part.
Transactions Have Their Own Retry Rules
Operations inside a transaction aren't individually retried. Instead, MongoDB labels transaction errors so the whole transaction can be retried:
TransientTransactionError: retry the entire transaction from the start.UnknownTransactionCommitResult: retry just the commit, which is itself safe to repeat.
The driver's callback API, session.withTransaction(), handles both labels and keeps retrying for up to 120 seconds. That's the main reason to use it instead of manual startTransaction() and commitTransaction() calls. MongoDB Transactions: When and How to Use Multi-Document ACID Transactions covers the details.
Testing Failure Handling with Fail Points
You don't need to kill servers to test retries. MongoDB has fail points, test hooks that make the server fail specific commands on purpose. They're available when the server runs with test commands enabled, so use them only on local or CI clusters:
mongod --replSet rs0 --dbpath ./data --setParameter enableTestCommands=1
Then configure the failCommand fail point to fail the next insert with a retryable error:
db.adminCommand({
configureFailPoint: "failCommand",
mode: { times: 1 },
data: {
failCommands: ["insert"],
errorCode: 91, // ShutdownInProgress
errorLabels: ["RetryableWriteError"],
},
});
Now run an insert from your application. The first attempt fails, the driver retries, and the insert succeeds. Your code never sees the error. Change times: 1 to times: 2, and both attempts fail, so the error reaches your code, which is exactly the scenario your application-level handling needs to cover.
You can also simulate a dropped connection with closeConnection: true:
db.adminCommand({
configureFailPoint: "failCommand",
mode: { times: 1 },
data: { failCommands: ["find"], closeConnection: true },
});
Turn the fail point off afterwards:
db.adminCommand({ configureFailPoint: "failCommand", mode: "off" });
Wiring this into integration tests is a cheap way to prove that your error handling works before a real election does it for you.
Application-Level Retries
The driver's single retry covers the common case, a quick election or one dropped connection. Some situations need more:
- An election that takes longer than the retry window.
- Non-retryable operations like
updateMany. - Errors from outside MongoDB in the same unit of work.
A small wrapper with exponential backoff handles these. The key is to retry only errors that are actually transient:
import {
MongoError,
MongoNetworkError,
MongoServerSelectionError,
} from "mongodb";
function isTransient(err) {
if (err instanceof MongoNetworkError) return true;
if (err instanceof MongoServerSelectionError) return true;
if (err instanceof MongoError) {
return (
err.hasErrorLabel("RetryableWriteError") ||
err.hasErrorLabel("TransientTransactionError")
);
}
return false;
}
export async function withRetry(fn, { attempts = 4, baseMs = 200 } = {}) {
for (let i = 1; ; i++) {
try {
return await fn();
} catch (err) {
if (i >= attempts || !isTransient(err)) throw err;
const delay = baseMs * 2 ** (i - 1) + Math.random() * baseMs;
console.warn(
`Transient MongoDB error, retry ${i} in ${Math.round(delay)}ms`,
err.message,
);
await new Promise((r) => setTimeout(r, delay));
}
}
}
// Only wrap operations that are idempotent
await withRetry(() =>
products.updateMany(
{ category: "lamps", "promo.fall2026": { $ne: true } },
{ $mul: { price: 0.9 }, $set: { "promo.fall2026": true } },
),
);
Two rules keep this safe. First, never retry validation errors, duplicate key errors, or authorization failures; they'll fail the same way every time. Second, only wrap operations you've made idempotent. Retrying a non-idempotent updateMany after an unknown outcome is exactly the duplicate-write bug you were trying to avoid. For a catalog of which error codes mean what, see MongoDB Error Handling: Common Error Codes and How to Fix Them.
Idempotency Keys for API Requests
Retries also happen above the database: a mobile client resends a request after a timeout, or a queue redelivers a message. Driver-level retries can't help here, because each request is a new operation. The standard solution is an idempotency key supplied by the client and enforced with a unique index:
await db
.collection("payments")
.createIndex({ idempotencyKey: 1 }, { unique: true });
async function createPayment(req) {
try {
await db.collection("payments").insertOne({
idempotencyKey: req.headers["idempotency-key"],
orderId: req.body.orderId,
amount: req.body.amount,
createdAt: new Date(),
});
return { status: 201 };
} catch (err) {
if (err.code === 11000) {
const existing = await db
.collection("payments")
.findOne({ idempotencyKey: req.headers["idempotency-key"] });
return { status: 200, body: existing }; // same result as the first request
}
throw err;
}
}
The first request creates the payment. Any repeat with the same key hits the unique index, and the handler returns the existing record instead of charging twice.
Timeouts and Failover Windows
Retries only help if your timeouts leave room for them. A typical election takes around 10 to 12 seconds with default settings. The settings that matter:
serverSelectionTimeoutMS(default 30,000): how long the driver waits for a suitable server, including during a retry. Lowering it below about 15 seconds risks failing operations during normal elections.socketTimeoutMSandconnectTimeoutMS: limits on individual network operations. Leave socket timeouts generous so long-running queries aren't cut off.timeoutMS: newer drivers support a client-side operation timeout that bounds the total time for an operation, including retries. Check your driver's documentation, since support and maturity vary by driver version.
On the application side, make sure your HTTP request timeout is longer than the time a retried operation might take, or users will see timeouts even though the database recovered.
Common Pitfalls
Disabling retries because of an old guide. Some older tutorials added retryWrites=false for compatibility with long-gone storage engines or tools. Remove it; there's no benefit on modern clusters.
Testing only on a standalone. Retryable writes silently do nothing on a standalone server, so tests pass without exercising the retry path. Use a single-node replica set in development and CI.
Assuming updateMany is retried. It isn't. Make multi-document updates idempotent, or break them into batches of updateOne or bulkWrite calls, which are retryable.
Retrying everything in a generic wrapper. A wrapper that retries any error with backoff turns a duplicate key error into five duplicate key errors and a slower failure. Check error types and labels.
Relying on retries for long cursors. Only the first batch of a read is retried. Long scans should checkpoint their position so they can restart after an error.
Forgetting idempotency at the API layer. Driver retries protect a single operation. Client retries, queue redeliveries, and double clicks need idempotency keys or natural unique constraints.
Conclusion
Retryable writes attach a session ID and transaction number to each single-document write so the server can apply a retried write at most once, even across failovers. Retryable reads give reads a second chance after transient errors. Both are on by default and handle the everyday failures of a replica set without any code. Around them, make multi-document writes idempotent, use withTransaction for transactions, add a narrow backoff wrapper for transient errors, and enforce idempotency keys for requests that can arrive twice.
For a concrete next step, spin up a single-node replica set with enableTestCommands=1, set the failCommand fail point with times: 1 on insert, and run your application's signup flow. Then set times: 2 and check that your code returns a clean, retry-friendly error instead of a stack trace.


