
Data Seeding and Fixtures for MongoDB Development Environments
Every team eventually hits the same wall. A new developer clones the repo, starts the app, and sees an empty dashboard. Or someone reports a bug that only happens for users with more than 50 orders, and nobody's local database has a user like that. The quick fix is usually "grab a dump from production", which works right up until someone's laptop with 400,000 customer email addresses on it gets stolen.
Seed data and fixtures solve this properly. Seed data is a realistic, reasonably sized dataset that fills a development database so the app looks and behaves like the real thing. Fixtures are small, precise, known datasets that tests and demos depend on. Both should be generated by code that lives in your repository, runs in seconds, and produces the same result every time.
This guide covers the difference between seeds and fixtures, writing an idempotent seed script with the Node.js driver, generating realistic fake data with Faker, keeping relationships consistent, loading JSON fixtures with mongoimport and mongorestore, seeding Docker containers automatically, and safely using production-shaped data.
Seeds vs. Fixtures
The terms get used interchangeably, but it helps to separate them:
| Seed data | Fixtures | |
|---|---|---|
| Purpose | Make dev environments usable | Give tests and demos known state |
| Size | Hundreds to tens of thousands of docs | A handful of docs |
| Content | Realistic, often randomly generated | Hand-written, precise |
| Stability | Can vary, ideally deterministic | Must be exactly the same every run |
| Used by | Developers, designers, QA | Automated tests, demo accounts |
A good setup usually has both: a seed script for local development and a fixtures/ folder that tests load. Some fixtures (like a known admin user) are also part of the seed so developers can log in.
Project Layout
A structure that scales well:
scripts/
seed.js # entry point: npm run seed
seed/
users.js # generators per collection
products.js
orders.js
fixtures/
users.json # hand-written fixtures for tests
products.json
Add scripts to package.json:
{
"scripts": {
"seed": "node scripts/seed.js",
"seed:reset": "node scripts/seed.js --reset"
}
}
Writing an Idempotent Seed Script
The most important property of a seed script is that you can run it again without breaking anything. Running npm run seed twice should not create duplicate users or fail with duplicate key errors. There are two approaches: reset (drop and recreate) or upsert (insert or update by a stable key).
Here's a seed script that supports both, using the official driver:
// scripts/seed.js
import { MongoClient } from "mongodb";
import { buildUsers } from "../seed/users.js";
import { buildProducts } from "../seed/products.js";
import { buildOrders } from "../seed/orders.js";
const uri = process.env.MONGODB_URI ?? "mongodb://localhost:27017";
const dbName = process.env.MONGODB_DB ?? "shop_dev";
const reset = process.argv.includes("--reset");
if (/prod/i.test(dbName) || process.env.NODE_ENV === "production") {
console.error("Refusing to seed a production database.");
process.exit(1);
}
const client = new MongoClient(uri);
try {
await client.connect();
const db = client.db(dbName);
if (reset) {
await db.dropDatabase();
console.log(`Dropped ${dbName}`);
}
await db.collection("users").createIndex({ email: 1 }, { unique: true });
await db.collection("products").createIndex({ sku: 1 }, { unique: true });
await db.collection("orders").createIndex({ userId: 1, createdAt: -1 });
const users = buildUsers(200);
const products = buildProducts(500);
const orders = buildOrders({ users, products, count: 2000 });
await upsertAll(db.collection("users"), users, "email");
await upsertAll(db.collection("products"), products, "sku");
await upsertAll(db.collection("orders"), orders, "_id");
console.log(
`Seeded ${users.length} users, ${products.length} products, ${orders.length} orders`,
);
} finally {
await client.close();
}
async function upsertAll(collection, docs, key) {
const ops = docs.map((doc) => ({
replaceOne: { filter: { [key]: doc[key] }, replacement: doc, upsert: true },
}));
for (let i = 0; i < ops.length; i += 1000) {
await collection.bulkWrite(ops.slice(i, i + 1000), { ordered: false });
}
}
A few details are worth calling out:
- The production guard. Seed scripts that drop databases must refuse to run against production. Checking the database name and
NODE_ENVis crude but catches the most common mistake. Better still, make sure your development credentials simply have no access to production. - Indexes are created in the seed. This keeps dev environments consistent with production and makes the unique keys the upsert relies on actually unique. In a mature project, running your migrations here instead is even better.
- Upserts by a natural key.
replaceOnewithupsert: truemakes reruns safe. Bulk writes in chunks of 1,000 keep it fast; see bulk write operations for the details.
Generating Realistic Data with Faker
Hand-writing 200 users isn't practical. @faker-js/faker generates plausible names, emails, addresses, prices, and dates:
npm install --save-dev @faker-js/faker
The key to useful seed data is determinism. If every run produces different random data, a developer can't say "look at user 17" and have a teammate see the same thing. Seed Faker's random generator with a fixed number:
// seed/users.js
import { faker } from "@faker-js/faker";
import { ObjectId } from "mongodb";
faker.seed(20260921);
export function buildUsers(count) {
const admin = {
_id: new ObjectId(fixedHex(1)),
email: "admin@example.test",
name: "Dev Admin",
role: "admin",
createdAt: new Date("2025-01-01T00:00:00Z"),
};
const users = Array.from({ length: count - 1 }, (_, i) => {
const firstName = faker.person.firstName();
const lastName = faker.person.lastName();
return {
_id: new ObjectId(fixedHex(i + 2)),
email: faker.internet
.email({ firstName, lastName, provider: "example.test" })
.toLowerCase(),
name: `${firstName} ${lastName}`,
role: faker.helpers.weightedArrayElement([
{ weight: 95, value: "customer" },
{ weight: 5, value: "staff" },
]),
address: {
street: faker.location.streetAddress(),
city: faker.location.city(),
country: faker.location.countryCode(),
},
createdAt: faker.date.between({ from: "2024-01-01", to: "2026-09-01" }),
};
});
return [admin, ...users];
}
export function fixedHex(n) {
return "65" + n.toString(16).padStart(22, "0");
}
Some choices here that pay off later:
- Fixed
_idvalues. Deterministic ObjectIds mean links like/users/650000000000000000000001work on every machine. Generating ObjectIds normally embeds the current time, so they'd differ per run. - The
.testdomain. It's reserved and will never deliver email, so a misconfigured mailer in dev can't email a real person. - Weighted distributions. Real data isn't uniform. Most users are customers, a few are staff. Mirroring realistic proportions surfaces UI and performance issues that uniform data hides.
- Fixed date ranges.
faker.date.recent()depends on the current date, which breaks determinism. Use explicit ranges.
Products follow the same pattern, with prices stored as integer cents so totals add up exactly:
// seed/products.js
import { faker } from "@faker-js/faker";
import { ObjectId } from "mongodb";
export function buildProducts(count) {
return Array.from({ length: count }, (_, i) => ({
_id: new ObjectId("66" + (i + 1).toString(16).padStart(22, "0")),
sku: `SKU-${String(i + 1).padStart(5, "0")}`,
name: faker.commerce.productName(),
category: faker.commerce.department(),
priceCents: faker.number.int({ min: 199, max: 49999 }),
stock: faker.number.int({ min: 0, max: 250 }),
}));
}
One caveat: Faker's output for a given seed can change between major versions of the library, so pin its version if you need exact reproducibility over time.
Keeping Relationships Consistent
Random data gets tricky when collections reference each other. An order must point to a user and products that actually exist, with a total that matches its line items. Build child collections from the parents you already generated:
// seed/orders.js
import { faker } from "@faker-js/faker";
import { ObjectId } from "mongodb";
export function buildOrders({ users, products, count }) {
const customers = users.filter((u) => u.role === "customer");
return Array.from({ length: count }, (_, i) => {
const user = faker.helpers.arrayElement(customers);
const picked = faker.helpers.arrayElements(products, { min: 1, max: 5 });
const items = picked.map((p) => {
const qty = faker.number.int({ min: 1, max: 3 });
return {
productId: p._id,
sku: p.sku,
name: p.name,
priceCents: p.priceCents,
qty,
};
});
const totalCents = items.reduce(
(sum, it) => sum + it.priceCents * it.qty,
0,
);
return {
_id: new ObjectId("70" + i.toString(16).padStart(22, "0")),
userId: user._id,
items,
totalCents,
status: faker.helpers.weightedArrayElement([
{ weight: 70, value: "delivered" },
{ weight: 15, value: "shipped" },
{ weight: 10, value: "pending" },
{ weight: 5, value: "cancelled" },
]),
createdAt: faker.date.between({ from: user.createdAt, to: "2026-09-15" }),
};
});
}
Notice that order dates are generated after the user's createdAt, and totals are computed rather than random. Inconsistencies like orders placed before the customer signed up won't crash anything, but they make dashboards and reports look wrong and erode trust in the dev environment.
Also include edge cases on purpose. Add a user with 150 orders (to test pagination), a product with a 200-character name (to test truncation), a user with no orders at all, and a product with a Unicode name. Random generation rarely produces these by chance.
export const edgeCaseProducts = [
{
sku: "EDGE-LONG",
name: "Ultra-Premium Handcrafted ".repeat(8).trim(),
priceCents: 129900,
},
{ sku: "EDGE-FREE", name: "Free Sample", priceCents: 0 },
{ sku: "EDGE-UNICODE", name: "Café Crème Brûlée Kit ☕", priceCents: 2450 },
];
Loading JSON Fixtures
For small, hand-written fixtures, JSON files are easy to review in pull requests. Use Extended JSON so types like ObjectId and dates survive the round trip:
[
{
"_id": { "$oid": "650000000000000000000001" },
"email": "admin@example.test",
"name": "Dev Admin",
"role": "admin",
"createdAt": { "$date": "2025-01-01T00:00:00Z" }
}
]
Load them in Node.js with the driver's EJSON parser (from the bson package, re-exported by mongodb):
import { readFile } from "node:fs/promises";
import { EJSON } from "bson";
export async function loadFixture(db, name) {
const raw = await readFile(
new URL(`../fixtures/${name}.json`, import.meta.url),
"utf8",
);
const docs = EJSON.parse(raw);
await db.collection(name).deleteMany({});
if (docs.length) await db.collection(name).insertMany(docs);
return docs;
}
Plain JSON.parse would leave { "$oid": ... } as a literal object, which is a classic source of "my fixture user can't be found by ID" bugs.
Using mongoimport and mongorestore
From the command line, mongoimport from the MongoDB Database Tools reads Extended JSON arrays directly:
mongoimport --uri "mongodb://localhost:27017/shop_dev" \
--collection users --file fixtures/users.json --jsonArray --drop
For larger datasets, a BSON dump is faster and preserves every type exactly. Generate it once from a seeded database and commit (or store) the archive:
mongodump --uri "mongodb://localhost:27017/shop_dev" --archive=seed/shop_dev.archive --gzip
mongorestore --uri "mongodb://localhost:27017" --archive=seed/shop_dev.archive --gzip --drop
mongorestore also recreates indexes from the dump. If your dump contains more than a few megabytes, keep it out of Git and fetch it from object storage in a setup script. See importing and exporting with mongoimport and mongoexport for more on these tools.
Seeding Docker Containers Automatically
The official mongo Docker image runs any .js or .sh files in /docker-entrypoint-initdb.d/ the first time the container starts with an empty data directory. That's a convenient hook for seeding:
# docker-compose.yml
services:
mongo:
image: mongo:8.0
ports:
- "27017:27017"
environment:
MONGO_INITDB_DATABASE: shop_dev
volumes:
- mongo-data:/data/db
- ./docker/mongo-init:/docker-entrypoint-initdb.d:ro
- ./fixtures:/fixtures:ro
volumes:
mongo-data:
#!/bin/bash
# docker/mongo-init/01-seed.sh
set -e
for file in /fixtures/*.json; do
collection=$(basename "$file" .json)
mongoimport --db shop_dev --collection "$collection" --file "$file" --jsonArray
done
.js files in the init directory run with mongosh against the MONGO_INITDB_DATABASE database, so you can also create indexes there with db.users.createIndex(...).
Remember that init scripts only run when the data volume is empty. To re-seed, remove the volume with docker compose down -v, or run your Node seed script against the running container instead. Running the Node script is often better because it's the same code path everywhere, whether a developer uses Docker, a local install, or an Atlas dev cluster.
Fixtures in Tests
Tests need fixtures that are small and explicit. The test should make it obvious what data exists, so a reader doesn't have to open three JSON files to understand an assertion. A lightweight factory function is often clearer than a shared fixture file:
let counter = 0;
export function makeUser(overrides = {}) {
counter += 1;
return {
email: `user${counter}@example.test`,
name: `Test User ${counter}`,
role: "customer",
createdAt: new Date("2026-01-01T00:00:00Z"),
...overrides,
};
}
// In a test:
await db
.collection("users")
.insertMany([
makeUser({ role: "admin" }),
makeUser(),
makeUser({ createdAt: new Date("2020-06-01") }),
]);
Use shared JSON fixtures for data that many tests depend on and that rarely changes (countries, currencies, plan definitions), and factories for everything else.
Working With Production-Shaped Data Safely
Sometimes you really do need data that looks like production: a performance issue that only appears at real cardinality, or a bug triggered by some odd historical document. Don't copy raw production data to laptops. Instead:
- Anonymize during export. Run an aggregation that replaces personal fields, then dump the result:
db.users.aggregate([
{ $sample: { size: 50000 } },
{
$set: {
email: { $concat: ["user_", { $toString: "$_id" }, "@example.test"] },
name: "Redacted User",
phone: "$$REMOVE",
"address.street": "1 Example Street",
},
},
{ $out: { db: "shop_sanitized", coll: "users" } },
]);
- Keep the shape, drop the substance. Field types, array lengths, and value distributions are what matter for reproducing bugs and performance problems, not the actual names.
- Restrict access. Sanitized datasets still belong in a shared, access-controlled location, not in Git.
Common Mistakes
Non-deterministic seeds. Without faker.seed() and fixed IDs, every run differs and "check user 17" means nothing. Seed the generator and derive IDs from indexes.
Seed scripts that aren't idempotent. Plain insertMany fails with duplicate key errors on the second run. Upsert by a stable key or drop and recreate explicitly with a --reset flag.
No production guard. A seed script that calls dropDatabase() must refuse to run against production. Check the environment and the database name, and keep production credentials off developer machines.
Parsing Extended JSON with JSON.parse. ObjectIds and dates become plain objects and strings. Use EJSON.parse or mongoimport.
Only happy-path data. Uniform random data rarely contains the edge cases that break UIs. Add them explicitly.
Conclusion
Good seed data makes a development environment feel real, and good fixtures make tests readable and reliable. Keep both in code, make seeds deterministic and idempotent, build related collections from their parents so references and totals stay consistent, use Extended JSON for fixture files, and never ship raw production data to development machines.
Start small: write a seed script that creates one known admin user and 50 generated records for your most important collection, add a production guard, and wire it into npm run seed. The next person who clones your repo will have a working app in under a minute.


