Type something to search...
Common MongoDB Mistakes Developers Make and How to Avoid Them

Common MongoDB Mistakes Developers Make and How to Avoid Them

MongoDB is easy to start with. You install a driver, call insertOne, and you have a working database in minutes. That low barrier is a strength, but it also means many teams ship to production before they've learned how MongoDB actually wants to be used. The problems don't show up on day one. They show up at 50,000 users, when a page that loaded in 40 ms now takes four seconds, or the Atlas bill doubles, or a document mysteriously hits 16 MB.

The encouraging part is that the same mistakes come up again and again, across teams and industries. Most of them are cheap to fix early and painful to fix late, and nearly all of them can be spotted with a few commands.

This guide walks through the most common MongoDB mistakes in schema design, indexing, querying, data types, and operations, with a concrete way to detect and fix each one.

Schema Design Mistakes

Designing Tables Instead of Documents

The single most common mistake is carrying a relational schema straight into MongoDB: a users collection, an addresses collection, a user_preferences collection, and a phone_numbers collection, all joined with $lookup on every request.

MongoDB's core design rule is data that is accessed together should be stored together. If you always display a user with their addresses and preferences, embed them:

{
  _id: ObjectId("66f1c2a9e4b0a1b2c3d4e5f6"),
  email: "ada@example.com",
  name: "Ada Lovelace",
  addresses: [
    { label: "home", city: "London", postcode: "W1 1AA", default: true }
  ],
  preferences: { newsletter: true, theme: "dark" }
}

One read, one document, no joins. Use references when the related data is large, shared by many parents, or updated independently at high frequency. Start by listing your application's most frequent queries, then design documents that answer each one in a single read. See one-to-many relationships in MongoDB for the full decision framework.

Unbounded Arrays

The opposite mistake is embedding things that grow forever: every comment on a post, every login event on a user, every reading from a sensor.

// This document will eventually break
db.posts.updateOne({ _id: postId }, { $push: { comments: newComment } });

Documents have a hard 16 MB limit, but you'll feel pain long before that. Every update rewrites a growing document, indexes on array fields get enormous (one index entry per element), and reading the post pulls every comment across the wire even when the page shows five.

How to detect it: find your largest documents.

db.posts.aggregate([
  {
    $project: {
      size: { $bsonSize: "$$ROOT" },
      comments: { $size: { $ifNull: ["$comments", []] } },
    },
  },
  { $sort: { size: -1 } },
  { $limit: 5 },
]);

How to fix it: move the unbounded data into its own collection with a reference to the parent, and optionally keep a small, bounded subset embedded (the subset pattern), such as the three most recent comments:

db.posts.updateOne(
  { _id: postId },
  {
    $push: { recentComments: { $each: [newComment], $slice: -3 } },
    $inc: { commentCount: 1 },
  },
);
db.comments.insertOne({ postId, ...newComment });

No Schema at All

"Schemaless" is a feature for iteration speed, not an invitation to let every code path write whatever it wants. Without any enforcement, you end up with createdAt stored as a Date in some documents, a string in others, and a Unix timestamp in a third group. Every query has to cope.

Add a $jsonSchema validator to important collections, even a loose one that only checks types of critical fields. It costs almost nothing and catches bugs at write time instead of at report time.

Indexing Mistakes

Missing Indexes

Without a suitable index, every query scans the entire collection. That's invisible with 1,000 documents and catastrophic with 10 million. The fix is to run explain() on your important queries and look for COLLSCAN:

db.orders.find({ customerId: 42, status: "open" }).explain("executionStats");

The numbers to compare are totalDocsExamined and nReturned. If the query examined 2,000,000 documents to return 12, you need an index:

db.orders.createIndex({ customerId: 1, status: 1 });

Turn on the profiler in development, or check the Atlas Performance Advisor in production, to find slow queries you didn't know about.

Compound Indexes in the Wrong Order

A compound index on { createdAt: 1, status: 1 } is far less useful for "open orders, newest first" than { status: 1, createdAt: -1 }. The guideline is ESR: Equality, Sort, Range. Put fields you match exactly first, then fields you sort on, then fields you filter by range:

// Query: status equals "open", sorted by createdAt, total greater than 100
db.orders.find({ status: "open", total: { $gt: 100 } }).sort({ createdAt: -1 });

// ESR-ordered index
db.orders.createIndex({ status: 1, createdAt: -1, total: 1 });

Too Many Indexes

The reverse problem: an index on every field "just in case". Each index slows down every write, consumes RAM, and competes for the WiredTiger cache. Check which indexes are actually used:

db.orders.aggregate([
  { $indexStats: {} },
  { $project: { name: 1, "accesses.ops": 1 } },
]);

Indexes with zero or near-zero ops over a representative period are candidates for removal. Hide them first with db.orders.hideIndex("name"), watch for regressions, and then drop them. Also look for redundant prefixes: if you have both { a: 1 } and { a: 1, b: 1 }, the first is usually unnecessary.

Query Mistakes

Fetching Whole Documents When You Need Two Fields

A list page that shows names and prices doesn't need each product's 40 KB description and 30-element review array. Use projections:

const products = await db
  .collection("products")
  .find(
    { category: "lamps" },
    { projection: { name: 1, priceCents: 1, thumbnail: 1 } },
  )
  .limit(24)
  .toArray();

This cuts network transfer and memory use in your app, and if the index covers all projected fields, MongoDB doesn't even need to read the documents.

Deep Pagination with skip()

skip(50000) still walks past 50,000 documents on every request. Page 1 is fast; page 2,000 is slow. For feeds, infinite scroll, and APIs, use range-based pagination on an indexed field:

const page = await db
  .collection("posts")
  .find(lastId ? { _id: { $lt: lastId } } : {})
  .sort({ _id: -1 })
  .limit(20)
  .toArray();

Details and trade-offs are in skip/limit vs. range-based pagination.

Unanchored and Case-Insensitive Regex for Search

{ name: /lamp/i } can't use an index efficiently, so it scans every index key or document. It's fine for admin tools on small collections, but it's a common cause of slow search boxes. Use a prefix regex (/^lamp/) for autocomplete on case-sensitive data, a collation-based index for case-insensitive exact matches, or a text index or Atlas Search for real search.

Counting Everything, Every Time

Showing "12,483,112 results" on every page load means countDocuments() scans the matching index range (or collection) on every request. For total collection size, estimatedDocumentCount() reads metadata and is instant. For filtered counts, consider caching, showing "10,000+", or maintaining a counter document updated with $inc.

Data Type Mistakes

Storing ObjectIds as Strings (Sometimes)

A classic bug: users._id is an ObjectId, but orders.userId was saved as a string from a request parameter. Now $lookup returns nothing and find({ userId: user._id }) matches nothing. No error, just empty results.

// Detect mixed types in a reference field
db.orders.aggregate([
  { $group: { _id: { $type: "$userId" }, count: { $sum: 1 } } },
]);
[
  { "_id": "objectId", "count": 181204 },
  { "_id": "string", "count": 3411 }
]

Convert at the API boundary, always, and fix existing data with an update pipeline:

db.orders.updateMany({ userId: { $type: "string" } }, [
  { $set: { userId: { $toObjectId: "$userId" } } },
]);

Dates as Strings

"2026-09-23" and "09/23/2026" in the same field make range queries and sorting unreliable, and even consistently formatted strings can't use date operators like $dateTrunc without conversion. Store real BSON dates in UTC and convert to local time only for display.

Money as Floating Point

0.1 + 0.2 is 0.30000000000000004 in MongoDB just as in JavaScript. For prices and balances, store integer minor units (cents) or use Decimal128. Floating-point rounding errors in financial totals are the kind of bug that gets discovered by an accountant, not a test.

Connection and Application Mistakes

Creating a New Client Per Request

This is the most damaging application-level mistake:

// Don't: opens a new pool (and new connections) for every request
app.get("/products", async (req, res) => {
  const client = new MongoClient(process.env.MONGODB_URI);
  await client.connect();
  const products = await client
    .db()
    .collection("products")
    .find()
    .limit(20)
    .toArray();
  await client.close();
  res.json(products);
});

Each MongoClient has its own connection pool, and establishing a connection (TCP, TLS, authentication) takes far longer than the query itself. Under load this exhausts the server's connection limit. Create one client at startup and reuse it everywhere:

// db.js
import { MongoClient } from "mongodb";
export const client = new MongoClient(process.env.MONGODB_URI, {
  maxPoolSize: 50,
});
export const db = client.db("shop");

In serverless environments, cache the client in module scope so warm invocations reuse it. See MongoDB connection pooling.

N+1 Queries

Looping over results and querying for each one is just as much a problem in MongoDB as in SQL:

// 1 query for orders + 1 query per order
for (const order of orders) {
  order.customer = await customers.findOne({ _id: order.customerId });
}

Batch it with $in, or let the server do it with $lookup:

const ids = [...new Set(orders.map((o) => o.customerId))];
const byId = new Map(
  (await customers.find({ _id: { $in: ids } }).toArray()).map((c) => [
    c._id.toString(),
    c,
  ]),
);
orders.forEach((o) => (o.customer = byId.get(o.customerId.toString())));

Read-Modify-Write Instead of Atomic Updates

Fetching a document, changing it in JavaScript, and saving it back loses updates when two requests do it at the same time:

// Race condition: two concurrent requests both read views = 10 and both write 11
const post = await posts.findOne({ _id: id });
await posts.updateOne({ _id: id }, { $set: { views: post.views + 1 } });

Use update operators that apply atomically on the server:

await posts.updateOne({ _id: id }, { $inc: { views: 1 } });

For conditional changes (only decrement stock if it's positive), put the condition in the filter: updateOne({ _id: id, stock: { $gt: 0 } }, { $inc: { stock: -1 } }), and check modifiedCount.

Using Transactions for Everything

Multi-document transactions are there when you need them, but single-document operations are already atomic. A well-designed schema needs transactions only occasionally. Wrapping every request in a transaction adds latency, increases the chance of write conflicts, and hides schema designs that should embed related data instead.

Operational Mistakes

Running a Single Node in Production

A standalone mongod has no failover, no redundancy, and doesn't support transactions or change streams. If it goes down, your app goes down, and if its disk dies, your data might go with it. Production should be a replica set of at least three members. On Atlas, every cluster (including M0) is already a replica set.

Exposing MongoDB to the Internet Without Auth

Thousands of databases have been wiped by ransomware bots that scan for open port 27017. Always enable authentication, bind to private interfaces (net.bindIp), use firewalls or VPC peering, and require TLS. On Atlas, avoid adding 0.0.0.0/0 to the IP access list except briefly for debugging.

Never Testing Restores

A backup you've never restored is a hope, not a backup. Schedule a periodic restore into a staging environment and verify the data. It's the only way to know your backup process actually works and how long recovery takes.

Ignoring the Working Set

MongoDB performs best when your frequently accessed data and indexes fit in RAM (the WiredTiger cache). When they don't, reads go to disk and latency jumps. Watch cache usage and page faults, and size your instances based on working set, not total data size.

Quick Diagnostic Checklist

Run through these on any MongoDB deployment you inherit:

CheckCommand or place to look
Slow queriesProfiler, or Atlas Performance Advisor
Collection scansexplain("executionStats") on top queries
Unused indexes$indexStats
Largest documents$bsonSize aggregation
Mixed field types$group on $type of key fields
Connection countdb.serverStatus().connections
Replica set healthrs.status()
Last restore testYour runbook (if it's blank, that's the finding)

Conclusion

Almost every MongoDB horror story traces back to a few repeatable mistakes: relational schemas forced into documents, unbounded arrays, missing or misordered indexes, inconsistent data types, and a new connection pool per request. None of them are hard to fix once you know what to look for, and all of them are much cheaper to fix early.

Pick the query your application runs most often, run it with explain("executionStats"), and compare totalDocsExamined to nReturned. If the ratio is much above 1, you've found your first fix.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading