Type something to search...
Data Seeding and Fixtures for MongoDB Development Environments

Data Seeding and Fixtures for MongoDB Development Environments

Every team eventually hits the same wall. A new developer clones the repo, starts the app, and sees an empty dashboard. Or someone reports a bug that only happens for users with more than 50 orders, and nobody's local database has a user like that. The quick fix is usually "grab a dump from production", which works right up until someone's laptop with 400,000 customer email addresses on it gets stolen.

Seed data and fixtures solve this properly. Seed data is a realistic, reasonably sized dataset that fills a development database so the app looks and behaves like the real thing. Fixtures are small, precise, known datasets that tests and demos depend on. Both should be generated by code that lives in your repository, runs in seconds, and produces the same result every time.

This guide covers the difference between seeds and fixtures, writing an idempotent seed script with the Node.js driver, generating realistic fake data with Faker, keeping relationships consistent, loading JSON fixtures with mongoimport and mongorestore, seeding Docker containers automatically, and safely using production-shaped data.

Seeds vs. Fixtures

The terms get used interchangeably, but it helps to separate them:

Seed dataFixtures
PurposeMake dev environments usableGive tests and demos known state
SizeHundreds to tens of thousands of docsA handful of docs
ContentRealistic, often randomly generatedHand-written, precise
StabilityCan vary, ideally deterministicMust be exactly the same every run
Used byDevelopers, designers, QAAutomated tests, demo accounts

A good setup usually has both: a seed script for local development and a fixtures/ folder that tests load. Some fixtures (like a known admin user) are also part of the seed so developers can log in.

Project Layout

A structure that scales well:

scripts/
  seed.js            # entry point: npm run seed
seed/
  users.js           # generators per collection
  products.js
  orders.js
fixtures/
  users.json         # hand-written fixtures for tests
  products.json

Add scripts to package.json:

{
  "scripts": {
    "seed": "node scripts/seed.js",
    "seed:reset": "node scripts/seed.js --reset"
  }
}

Writing an Idempotent Seed Script

The most important property of a seed script is that you can run it again without breaking anything. Running npm run seed twice should not create duplicate users or fail with duplicate key errors. There are two approaches: reset (drop and recreate) or upsert (insert or update by a stable key).

Here's a seed script that supports both, using the official driver:

// scripts/seed.js
import { MongoClient } from "mongodb";
import { buildUsers } from "../seed/users.js";
import { buildProducts } from "../seed/products.js";
import { buildOrders } from "../seed/orders.js";

const uri = process.env.MONGODB_URI ?? "mongodb://localhost:27017";
const dbName = process.env.MONGODB_DB ?? "shop_dev";
const reset = process.argv.includes("--reset");

if (/prod/i.test(dbName) || process.env.NODE_ENV === "production") {
  console.error("Refusing to seed a production database.");
  process.exit(1);
}

const client = new MongoClient(uri);

try {
  await client.connect();
  const db = client.db(dbName);

  if (reset) {
    await db.dropDatabase();
    console.log(`Dropped ${dbName}`);
  }

  await db.collection("users").createIndex({ email: 1 }, { unique: true });
  await db.collection("products").createIndex({ sku: 1 }, { unique: true });
  await db.collection("orders").createIndex({ userId: 1, createdAt: -1 });

  const users = buildUsers(200);
  const products = buildProducts(500);
  const orders = buildOrders({ users, products, count: 2000 });

  await upsertAll(db.collection("users"), users, "email");
  await upsertAll(db.collection("products"), products, "sku");
  await upsertAll(db.collection("orders"), orders, "_id");

  console.log(
    `Seeded ${users.length} users, ${products.length} products, ${orders.length} orders`,
  );
} finally {
  await client.close();
}

async function upsertAll(collection, docs, key) {
  const ops = docs.map((doc) => ({
    replaceOne: { filter: { [key]: doc[key] }, replacement: doc, upsert: true },
  }));
  for (let i = 0; i < ops.length; i += 1000) {
    await collection.bulkWrite(ops.slice(i, i + 1000), { ordered: false });
  }
}

A few details are worth calling out:

  • The production guard. Seed scripts that drop databases must refuse to run against production. Checking the database name and NODE_ENV is crude but catches the most common mistake. Better still, make sure your development credentials simply have no access to production.
  • Indexes are created in the seed. This keeps dev environments consistent with production and makes the unique keys the upsert relies on actually unique. In a mature project, running your migrations here instead is even better.
  • Upserts by a natural key. replaceOne with upsert: true makes reruns safe. Bulk writes in chunks of 1,000 keep it fast; see bulk write operations for the details.

Generating Realistic Data with Faker

Hand-writing 200 users isn't practical. @faker-js/faker generates plausible names, emails, addresses, prices, and dates:

npm install --save-dev @faker-js/faker

The key to useful seed data is determinism. If every run produces different random data, a developer can't say "look at user 17" and have a teammate see the same thing. Seed Faker's random generator with a fixed number:

// seed/users.js
import { faker } from "@faker-js/faker";
import { ObjectId } from "mongodb";

faker.seed(20260921);

export function buildUsers(count) {
  const admin = {
    _id: new ObjectId(fixedHex(1)),
    email: "admin@example.test",
    name: "Dev Admin",
    role: "admin",
    createdAt: new Date("2025-01-01T00:00:00Z"),
  };

  const users = Array.from({ length: count - 1 }, (_, i) => {
    const firstName = faker.person.firstName();
    const lastName = faker.person.lastName();
    return {
      _id: new ObjectId(fixedHex(i + 2)),
      email: faker.internet
        .email({ firstName, lastName, provider: "example.test" })
        .toLowerCase(),
      name: `${firstName} ${lastName}`,
      role: faker.helpers.weightedArrayElement([
        { weight: 95, value: "customer" },
        { weight: 5, value: "staff" },
      ]),
      address: {
        street: faker.location.streetAddress(),
        city: faker.location.city(),
        country: faker.location.countryCode(),
      },
      createdAt: faker.date.between({ from: "2024-01-01", to: "2026-09-01" }),
    };
  });

  return [admin, ...users];
}

export function fixedHex(n) {
  return "65" + n.toString(16).padStart(22, "0");
}

Some choices here that pay off later:

  • Fixed _id values. Deterministic ObjectIds mean links like /users/650000000000000000000001 work on every machine. Generating ObjectIds normally embeds the current time, so they'd differ per run.
  • The .test domain. It's reserved and will never deliver email, so a misconfigured mailer in dev can't email a real person.
  • Weighted distributions. Real data isn't uniform. Most users are customers, a few are staff. Mirroring realistic proportions surfaces UI and performance issues that uniform data hides.
  • Fixed date ranges. faker.date.recent() depends on the current date, which breaks determinism. Use explicit ranges.

Products follow the same pattern, with prices stored as integer cents so totals add up exactly:

// seed/products.js
import { faker } from "@faker-js/faker";
import { ObjectId } from "mongodb";

export function buildProducts(count) {
  return Array.from({ length: count }, (_, i) => ({
    _id: new ObjectId("66" + (i + 1).toString(16).padStart(22, "0")),
    sku: `SKU-${String(i + 1).padStart(5, "0")}`,
    name: faker.commerce.productName(),
    category: faker.commerce.department(),
    priceCents: faker.number.int({ min: 199, max: 49999 }),
    stock: faker.number.int({ min: 0, max: 250 }),
  }));
}

One caveat: Faker's output for a given seed can change between major versions of the library, so pin its version if you need exact reproducibility over time.

Keeping Relationships Consistent

Random data gets tricky when collections reference each other. An order must point to a user and products that actually exist, with a total that matches its line items. Build child collections from the parents you already generated:

// seed/orders.js
import { faker } from "@faker-js/faker";
import { ObjectId } from "mongodb";

export function buildOrders({ users, products, count }) {
  const customers = users.filter((u) => u.role === "customer");

  return Array.from({ length: count }, (_, i) => {
    const user = faker.helpers.arrayElement(customers);
    const picked = faker.helpers.arrayElements(products, { min: 1, max: 5 });
    const items = picked.map((p) => {
      const qty = faker.number.int({ min: 1, max: 3 });
      return {
        productId: p._id,
        sku: p.sku,
        name: p.name,
        priceCents: p.priceCents,
        qty,
      };
    });
    const totalCents = items.reduce(
      (sum, it) => sum + it.priceCents * it.qty,
      0,
    );

    return {
      _id: new ObjectId("70" + i.toString(16).padStart(22, "0")),
      userId: user._id,
      items,
      totalCents,
      status: faker.helpers.weightedArrayElement([
        { weight: 70, value: "delivered" },
        { weight: 15, value: "shipped" },
        { weight: 10, value: "pending" },
        { weight: 5, value: "cancelled" },
      ]),
      createdAt: faker.date.between({ from: user.createdAt, to: "2026-09-15" }),
    };
  });
}

Notice that order dates are generated after the user's createdAt, and totals are computed rather than random. Inconsistencies like orders placed before the customer signed up won't crash anything, but they make dashboards and reports look wrong and erode trust in the dev environment.

Also include edge cases on purpose. Add a user with 150 orders (to test pagination), a product with a 200-character name (to test truncation), a user with no orders at all, and a product with a Unicode name. Random generation rarely produces these by chance.

export const edgeCaseProducts = [
  {
    sku: "EDGE-LONG",
    name: "Ultra-Premium Handcrafted ".repeat(8).trim(),
    priceCents: 129900,
  },
  { sku: "EDGE-FREE", name: "Free Sample", priceCents: 0 },
  { sku: "EDGE-UNICODE", name: "Café Crème Brûlée Kit ☕", priceCents: 2450 },
];

Loading JSON Fixtures

For small, hand-written fixtures, JSON files are easy to review in pull requests. Use Extended JSON so types like ObjectId and dates survive the round trip:

[
  {
    "_id": { "$oid": "650000000000000000000001" },
    "email": "admin@example.test",
    "name": "Dev Admin",
    "role": "admin",
    "createdAt": { "$date": "2025-01-01T00:00:00Z" }
  }
]

Load them in Node.js with the driver's EJSON parser (from the bson package, re-exported by mongodb):

import { readFile } from "node:fs/promises";
import { EJSON } from "bson";

export async function loadFixture(db, name) {
  const raw = await readFile(
    new URL(`../fixtures/${name}.json`, import.meta.url),
    "utf8",
  );
  const docs = EJSON.parse(raw);
  await db.collection(name).deleteMany({});
  if (docs.length) await db.collection(name).insertMany(docs);
  return docs;
}

Plain JSON.parse would leave { "$oid": ... } as a literal object, which is a classic source of "my fixture user can't be found by ID" bugs.

Using mongoimport and mongorestore

From the command line, mongoimport from the MongoDB Database Tools reads Extended JSON arrays directly:

mongoimport --uri "mongodb://localhost:27017/shop_dev" \
  --collection users --file fixtures/users.json --jsonArray --drop

For larger datasets, a BSON dump is faster and preserves every type exactly. Generate it once from a seeded database and commit (or store) the archive:

mongodump --uri "mongodb://localhost:27017/shop_dev" --archive=seed/shop_dev.archive --gzip
mongorestore --uri "mongodb://localhost:27017" --archive=seed/shop_dev.archive --gzip --drop

mongorestore also recreates indexes from the dump. If your dump contains more than a few megabytes, keep it out of Git and fetch it from object storage in a setup script. See importing and exporting with mongoimport and mongoexport for more on these tools.

Seeding Docker Containers Automatically

The official mongo Docker image runs any .js or .sh files in /docker-entrypoint-initdb.d/ the first time the container starts with an empty data directory. That's a convenient hook for seeding:

# docker-compose.yml
services:
  mongo:
    image: mongo:8.0
    ports:
      - "27017:27017"
    environment:
      MONGO_INITDB_DATABASE: shop_dev
    volumes:
      - mongo-data:/data/db
      - ./docker/mongo-init:/docker-entrypoint-initdb.d:ro
      - ./fixtures:/fixtures:ro

volumes:
  mongo-data:
#!/bin/bash
# docker/mongo-init/01-seed.sh
set -e
for file in /fixtures/*.json; do
  collection=$(basename "$file" .json)
  mongoimport --db shop_dev --collection "$collection" --file "$file" --jsonArray
done

.js files in the init directory run with mongosh against the MONGO_INITDB_DATABASE database, so you can also create indexes there with db.users.createIndex(...).

Remember that init scripts only run when the data volume is empty. To re-seed, remove the volume with docker compose down -v, or run your Node seed script against the running container instead. Running the Node script is often better because it's the same code path everywhere, whether a developer uses Docker, a local install, or an Atlas dev cluster.

Fixtures in Tests

Tests need fixtures that are small and explicit. The test should make it obvious what data exists, so a reader doesn't have to open three JSON files to understand an assertion. A lightweight factory function is often clearer than a shared fixture file:

let counter = 0;

export function makeUser(overrides = {}) {
  counter += 1;
  return {
    email: `user${counter}@example.test`,
    name: `Test User ${counter}`,
    role: "customer",
    createdAt: new Date("2026-01-01T00:00:00Z"),
    ...overrides,
  };
}

// In a test:
await db
  .collection("users")
  .insertMany([
    makeUser({ role: "admin" }),
    makeUser(),
    makeUser({ createdAt: new Date("2020-06-01") }),
  ]);

Use shared JSON fixtures for data that many tests depend on and that rarely changes (countries, currencies, plan definitions), and factories for everything else.

Working With Production-Shaped Data Safely

Sometimes you really do need data that looks like production: a performance issue that only appears at real cardinality, or a bug triggered by some odd historical document. Don't copy raw production data to laptops. Instead:

  • Anonymize during export. Run an aggregation that replaces personal fields, then dump the result:
db.users.aggregate([
  { $sample: { size: 50000 } },
  {
    $set: {
      email: { $concat: ["user_", { $toString: "$_id" }, "@example.test"] },
      name: "Redacted User",
      phone: "$$REMOVE",
      "address.street": "1 Example Street",
    },
  },
  { $out: { db: "shop_sanitized", coll: "users" } },
]);
  • Keep the shape, drop the substance. Field types, array lengths, and value distributions are what matter for reproducing bugs and performance problems, not the actual names.
  • Restrict access. Sanitized datasets still belong in a shared, access-controlled location, not in Git.

Common Mistakes

Non-deterministic seeds. Without faker.seed() and fixed IDs, every run differs and "check user 17" means nothing. Seed the generator and derive IDs from indexes.

Seed scripts that aren't idempotent. Plain insertMany fails with duplicate key errors on the second run. Upsert by a stable key or drop and recreate explicitly with a --reset flag.

No production guard. A seed script that calls dropDatabase() must refuse to run against production. Check the environment and the database name, and keep production credentials off developer machines.

Parsing Extended JSON with JSON.parse. ObjectIds and dates become plain objects and strings. Use EJSON.parse or mongoimport.

Only happy-path data. Uniform random data rarely contains the edge cases that break UIs. Add them explicitly.

Conclusion

Good seed data makes a development environment feel real, and good fixtures make tests readable and reliable. Keep both in code, make seeds deterministic and idempotent, build related collections from their parents so references and totals stay consistent, use Extended JSON for fixture files, and never ship raw production data to development machines.

Start small: write a seed script that creates one known admin user and 50 generated records for your most important collection, add a production guard, and wire it into npm run seed. The next person who clones your repo will have a working app in under a minute.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading