Type something to search...
Bulk Write Operations in MongoDB for Faster Data Imports

Bulk Write Operations in MongoDB for Faster Data Imports

You've written an import script. It reads a file of product records, loops over them, and calls insertOne for each. With 500 test records it finishes instantly. With the real file, 2 million records, it's still running after an hour, and the database barely looks busy. The bottleneck isn't MongoDB's ability to write. It's the round trips: every insertOne sends a request over the network, waits for the server to acknowledge it, and only then sends the next.

Bulk write operations fix this by sending many operations in a single request. Instead of two million round trips, you make a few hundred, and the server processes each batch efficiently. The difference is routinely one or two orders of magnitude. Bulk writes also let you mix inserts, updates, upserts, and deletes in a single call, which makes them the backbone of sync jobs and ETL pipelines, not just one-off imports.

This guide covers insertMany and bulkWrite, ordered versus unordered execution, how drivers split batches, how to stream a large file into MongoDB in chunks, how to handle partial failures, and the tuning options that matter for speed.

Why One-at-a-Time Writes Are Slow

Every individual write costs at least one network round trip. If your application server and database are in the same region, that's typically somewhere around half a millisecond to a few milliseconds. Multiply by two million and you're looking at somewhere between 15 minutes and over an hour of pure waiting, before counting the server's own work.

// Slow: one round trip per document
for (const product of products) {
  await db.collection("products").insertOne(product);
}

Batching changes the math. Send 1,000 documents per request and those two million writes take 2,000 round trips. The server still has to write every document and update every index, but it does so in a tight loop without waiting on the network between each one.

insertMany: The Simple Case

When every operation is an insert, insertMany is the easiest tool:

const result = await db.collection("products").insertMany([
  { sku: "LMP-204", name: "Brass Desk Lamp", price: 89.99 },
  { sku: "LMP-205", name: "Walnut Floor Lamp", price: 149.0 },
  { sku: "BLB-011", name: "Warm LED Bulb", price: 4.25 },
]);

console.log(result.insertedCount); // 3
console.log(result.insertedIds); // { '0': ObjectId(...), '1': ObjectId(...), '2': ObjectId(...) }

In mongosh the syntax is the same, and in PyMongo it's insert_many:

result = db.products.insert_many(products, ordered=False)
print(len(result.inserted_ids))

If documents don't have an _id, the driver generates ObjectIds client-side before sending, which is why insertedIds is available immediately.

bulkWrite: Mixed Operations in One Call

bulkWrite accepts an array of operation descriptions, each one of insertOne, updateOne, updateMany, replaceOne, deleteOne, or deleteMany:

const result = await db.collection("products").bulkWrite([
  {
    insertOne: { document: { sku: "LMP-206", name: "Arc Lamp", price: 219.0 } },
  },
  {
    updateOne: {
      filter: { sku: "LMP-204" },
      update: { $set: { price: 84.99, updatedAt: new Date() } },
    },
  },
  {
    updateOne: {
      filter: { sku: "BLB-012" },
      update: { $set: { name: "Cool LED Bulb", price: 4.25 } },
      upsert: true,
    },
  },
  { deleteOne: { filter: { sku: "LMP-099" } } },
  {
    replaceOne: {
      filter: { sku: "LMP-205" },
      replacement: { sku: "LMP-205", name: "Walnut Floor Lamp", price: 139.0 },
    },
  },
]);

The result summarizes what happened across every operation:

{
  "insertedCount": 1,
  "matchedCount": 2,
  "modifiedCount": 2,
  "deletedCount": 1,
  "upsertedCount": 1,
  "upsertedIds": { "2": "66ea91c4d1b2e3f405162738" }
}

The keys in upsertedIds are the positions of the operations in your array, which lets you map generated IDs back to input records.

The same thing in PyMongo uses operation classes:

from datetime import datetime, timezone
from pymongo import InsertOne, UpdateOne, DeleteOne, ReplaceOne

ops = [
    InsertOne({"sku": "LMP-206", "name": "Arc Lamp", "price": 219.0}),
    UpdateOne({"sku": "LMP-204"}, {"$set": {"price": 84.99, "updatedAt": datetime.now(timezone.utc)}}),
    UpdateOne({"sku": "BLB-012"}, {"$set": {"name": "Cool LED Bulb", "price": 4.25}}, upsert=True),
    DeleteOne({"sku": "LMP-099"}),
    ReplaceOne({"sku": "LMP-205"}, {"sku": "LMP-205", "name": "Walnut Floor Lamp", "price": 139.0}),
]

result = db.products.bulk_write(ops, ordered=False)
print(result.bulk_api_result)

Ordered vs. Unordered

Every bulk operation is either ordered (the default) or unordered, and the choice affects both correctness and speed.

With ordered: true, the server executes operations in sequence and stops at the first error. Everything before the error is applied; everything after is never attempted. Use this when later operations depend on earlier ones, such as inserting a document and then updating it in the same batch.

With ordered: false, the server may execute operations in any order, and an error in one operation doesn't stop the others. At the end, you get a report of every failure. This is almost always what you want for imports, because one bad record shouldn't abort the other 999 in the batch.

Unordered writes can also be faster, particularly on sharded clusters, where the driver and mongos can send operations to different shards in parallel instead of preserving a strict sequence.

await db.collection("products").bulkWrite(ops, { ordered: false });

Neither mode is a transaction. Even with ordered writes, a failure halfway through leaves the first half applied. If you need all-or-nothing semantics, wrap the bulk write in a transaction, keeping in mind that transactions have their own size and time limits and are a poor fit for huge imports.

How Drivers Split Large Batches

You can pass insertMany an array of 500,000 documents, and it will work. The driver automatically splits it into multiple server commands based on two limits the server advertises:

  • maxWriteBatchSize: the maximum number of operations per command, currently 100,000.
  • maxMessageSizeBytes: the maximum size of a single wire protocol message, 48 MB.

Individual documents are separately capped at 16 MB each, which is the BSON document limit.

So why chunk your input yourself? Because the array has to exist in memory. Loading two million parsed documents into one JavaScript array can easily consume gigabytes of heap. For large imports, stream the source and flush a batch whenever it reaches a fixed size. Batches of 500 to 5,000 documents are a sensible range; beyond that the gains flatten out while memory use and retry costs keep growing.

Streaming a Large File Import

Here's a complete Node.js import that reads newline-delimited JSON (one document per line), writes in unordered batches, tolerates duplicates, and reports progress:

import { createReadStream } from "node:fs";
import { createInterface } from "node:readline";
import { MongoClient, MongoBulkWriteError } from "mongodb";

const BATCH_SIZE = 1000;

async function flush(collection, batch, stats) {
  if (batch.length === 0) return;
  try {
    const res = await collection.insertMany(batch, { ordered: false });
    stats.inserted += res.insertedCount;
  } catch (err) {
    if (!(err instanceof MongoBulkWriteError)) throw err;
    const errors = [].concat(err.writeErrors ?? []);
    const fatal = errors.filter((e) => e.code !== 11000);
    stats.inserted += err.result.insertedCount;
    stats.duplicates += errors.length - fatal.length;
    stats.failed += fatal.length;
    for (const e of fatal.slice(0, 5)) console.error("Write error:", e.errmsg);
  }
}

async function importFile(path, uri) {
  const client = new MongoClient(uri);
  await client.connect();
  const products = client.db("shop").collection("products");

  const stats = { inserted: 0, duplicates: 0, failed: 0, invalid: 0 };
  const lines = createInterface({
    input: createReadStream(path),
    crlfDelay: Infinity,
  });
  let batch = [];
  let batches = 0;

  try {
    for await (const line of lines) {
      if (!line.trim()) continue;
      try {
        batch.push(JSON.parse(line));
      } catch {
        stats.invalid++;
        continue;
      }
      if (batch.length >= BATCH_SIZE) {
        await flush(products, batch, stats);
        batch = [];
        if (++batches % 50 === 0) console.log(stats);
      }
    }
    await flush(products, batch, stats);
  } finally {
    await client.close();
  }

  console.log("Done:", stats);
}

await importFile("./products.ndjson", process.env.MONGODB_URI);
{ inserted: 50000, duplicates: 0, failed: 0, invalid: 0 }
{ inserted: 99988, duplicates: 12, failed: 0, invalid: 1 }
...
Done: { inserted: 1998412, duplicates: 1571, failed: 3, invalid: 14 }

A few design choices are worth pointing out. The for await loop naturally applies backpressure: the file isn't read faster than MongoDB can write. Duplicate key errors (code 11000) are counted, not treated as fatal, so reruns after a crash are safe with a unique index on sku. And parse errors are counted separately so bad input lines don't masquerade as database failures. For more on interpreting duplicate errors, see handling duplicate key errors.

Idempotent Sync with Upserts

Imports that run repeatedly (a nightly feed from a supplier, a sync from another system) shouldn't insert blindly. Convert each record into an upsert keyed on a natural identifier, so running the job twice produces the same result:

function toUpsert(record) {
  return {
    updateOne: {
      filter: { sku: record.sku },
      update: {
        $set: {
          name: record.name,
          price: record.price,
          stock: record.stock,
          syncedAt: new Date(),
        },
        $setOnInsert: { createdAt: new Date() },
      },
      upsert: true,
    },
  };
}

const res = await db
  .collection("products")
  .bulkWrite(records.map(toUpsert), { ordered: false });
console.log({ updated: res.modifiedCount, created: res.upsertedCount });

$setOnInsert sets fields only when the upsert creates a new document, so createdAt reflects the first import rather than the latest. Make sure the filter field (sku here) has a unique index: without it, each upsert must scan the collection to find its match, and concurrent upserts can create duplicates.

To remove records that disappeared from the feed, stamp each sync with a run ID and delete the stragglers at the end with a single deleteMany({ syncRunId: { $ne: currentRunId } }). Scope that delete carefully, since a truncated feed would otherwise wipe out valid data. A sanity check like "refuse to delete more than 5 percent of the collection" is cheap insurance.

Mongoose Bulk Operations

Mongoose offers both Model.insertMany() and Model.bulkWrite(). They behave differently from save() in ways that matter for imports:

// Validates and casts every document against the schema
await Product.insertMany(docs, { ordered: false });

// Casts filters and updates, but does not run document middleware like pre('save')
await Product.bulkWrite(
  docs.map((d) => ({
    updateOne: { filter: { sku: d.sku }, update: { $set: d }, upsert: true },
  })),
  { ordered: false },
);

insertMany runs validation, which is valuable for untrusted input but adds CPU cost; save() hooks don't fire for either method. If your schema relies on pre("save") middleware to compute fields, compute them in the import code instead. For the fastest possible path with already-trusted data, you can drop down to the native collection with Product.collection.insertMany(...), bypassing Mongoose entirely.

Bulk Writes Across Collections

Starting with MongoDB 8.0, the server supports a client-level bulk write that can target multiple collections in a single request. Recent versions of the official drivers expose it; in the Node.js driver it's client.bulkWrite(), where each operation names its namespace:

const result = await client.bulkWrite([
  {
    name: "insertOne",
    namespace: "shop.orders",
    document: { orderId: "ORD-1", total: 42 },
  },
  {
    name: "updateOne",
    namespace: "shop.inventory",
    filter: { sku: "LMP-204" },
    update: { $inc: { reserved: 1 } },
  },
]);

This is useful when an import fans out into several collections, since it saves one round trip per collection per batch. It requires an 8.0 or newer server and a driver version that supports it, so check both before relying on it. For most single-collection imports, collection.bulkWrite() remains the tool.

Tuning for Speed

Once you're batching, a few more levers can make a big difference.

Write concern. The default write concern for replica sets is w: "majority", which waits for a majority of members to acknowledge each batch. That's the right default for application data. For a bulk load you can re-run if it fails, w: 1 reduces latency per batch. Never use w: 0 (unacknowledged) for imports: you won't learn about errors at all.

await collection.insertMany(batch, { ordered: false, writeConcern: { w: 1 } });

Indexes. Every index must be updated for every inserted document. For an initial load into an empty collection, it's often faster to insert everything first and then build secondary indexes once, since an index build over existing data is more efficient than millions of incremental updates. Keep any unique index you depend on for deduplication in place, though.

Parallelism. A single stream of sequential batches may not saturate a large cluster. Running a few batches concurrently (two to eight is a reasonable starting range) can raise throughput. Beyond that, you mostly add contention. Measure rather than guess, and watch the cluster's CPU, disk, and replication lag while you do.

Document size and shape. Smaller documents write faster. Drop fields you don't need, and convert types during the import (dates as real Dates, numbers as numbers) so you don't need a second pass later.

Use the right tool. If the data is already in JSON or CSV files and needs little transformation, mongoimport does all of the above for you with a single command. The post on importing and exporting with mongoimport and mongoexport covers it. Write your own bulk loader when you need custom validation, transformation, or upsert logic.

Common Pitfalls

Using ordered writes for imports. One bad record halts the rest of the batch. Default to ordered: false unless operations depend on each other.

Loading the entire dataset into memory. The driver splits large arrays for you, but it can't stream a file you've already parsed into one giant array. Read and flush in chunks.

Ignoring partial failures. A bulk write error still means most operations may have succeeded. Inspect writeErrors and the result counts, and log a sample of failures.

Upserting without a unique index on the filter field. Every upsert becomes a collection scan, and concurrent upserts can create duplicates.

Treating bulk writes as transactions. Neither ordered nor unordered bulk writes are atomic as a whole. Design imports so that re-running them is safe.

Hammering production during business hours. A full-speed import competes with application traffic for cache and disk. Throttle the batch rate or schedule large loads off-peak, and keep an eye on replication lag.

Conclusion

Bulk writes are the single most effective change you can make to a slow import. insertMany handles plain inserts, bulkWrite handles mixed operations and idempotent upserts, and unordered execution keeps one bad record from blocking thousands of good ones. Stream your input in batches of around a thousand, handle partial failures explicitly, keep a unique index on your natural key so re-runs are safe, and tune write concern and indexes only after the basics are in place.

Take your slowest data-loading script, replace the per-document writes with batches of 1,000 using insertMany or bulkWrite with ordered: false, and time it before and after. The difference will make the case for you.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading