Type something to search...
MongoDB Interview Questions and Answers for Developers

MongoDB Interview Questions and Answers for Developers

MongoDB interviews tend to follow a pattern. They start with fundamentals to check you've actually used it (documents, BSON, _id), move into querying and indexing to see whether you can make it fast, and finish with data modeling and scaling questions that reveal whether you understand the trade-offs. Senior interviews lean heavily on that last part: there's rarely one right answer, and the interviewer wants to hear you reason about access patterns, consistency, and failure modes.

Memorizing definitions won't get you far. What does work is understanding why MongoDB behaves the way it does, so you can answer the follow-up question you didn't prepare for. Each answer below includes the core fact plus the reasoning an interviewer is listening for.

This guide covers questions in seven areas: fundamentals, querying, indexing, aggregation, data modeling, replication and sharding, and transactions and consistency, followed by a few scenario questions and tips for answering them well.

Fundamentals

What is MongoDB, and how is it different from a relational database?

MongoDB is a document database. Data is stored as documents (JSON-like objects) grouped into collections, rather than as rows in tables. Documents in the same collection can have different fields, and a single document can contain nested objects and arrays.

The practical differences an interviewer wants to hear: related data is often embedded in one document instead of spread across tables and joined; the schema is flexible by default (though you can enforce one with validation); and MongoDB was designed for horizontal scaling through sharding. It still supports secondary indexes, rich queries, aggregations, joins via $lookup, and multi-document ACID transactions.

What is BSON, and why doesn't MongoDB just store JSON?

BSON (Binary JSON) is the binary format MongoDB uses to store documents and send them over the wire. It exists for two reasons. First, it supports more types than JSON: dates, 32- and 64-bit integers, Decimal128, ObjectId, binary data, and more. JSON only has "number" and "string", so you'd lose type information. Second, it's designed for fast traversal: fields are length-prefixed, so the server can skip over parts of a document without parsing them.

A good follow-up point: the maximum BSON document size is 16 MB, which is a guardrail against unbounded document growth.

What is an ObjectId?

The default type for _id. It's a 12-byte value made of a 4-byte timestamp (seconds since the Unix epoch), a 5-byte random value unique to the machine and process, and a 3-byte incrementing counter. That structure means ObjectIds can be generated by the client without coordination, are roughly sortable by creation time, and let you extract the creation time with ObjectId.getTimestamp().

ObjectId("66e5382a1f2b3c4d5e6f7a8b").getTimestamp();
// ISODate("2024-09-14T07:15:54.000Z")

Is MongoDB schemaless?

"Schema-flexible" is more accurate. The database doesn't require documents in a collection to share a structure, but your application always has an implicit schema. You can enforce one with $jsonSchema validators at the collection level. Strong answers mention that flexibility is useful for iteration and polymorphic data, but that production systems usually add validation for critical fields.

Querying

What's the difference between find() and aggregate()?

find() filters documents, optionally projecting, sorting, skipping, and limiting them. aggregate() runs a pipeline of stages that can filter, reshape, group, join, compute new fields, and write results to another collection. Anything find() does, an aggregation can do, but find() is simpler and the right choice for straightforward retrieval.

How do you query inside arrays, and what does $elemMatch do?

A query on an array field matches if any element matches: { tags: "mongodb" } finds documents whose tags array contains "mongodb". The subtlety comes with multiple conditions on arrays of objects:

// Matches if ANY element has qty > 10 and ANY (possibly different) element has size "M"
db.orders.find({ "items.qty": { $gt: 10 }, "items.size": "M" });

// Matches only if a SINGLE element satisfies both conditions
db.orders.find({ items: { $elemMatch: { qty: { $gt: 10 }, size: "M" } } });

Interviewers ask this because the first form is a common bug.

What's the difference between $set and replacing a document?

updateOne(filter, { $set: { status: "paid" } }) modifies only the named fields. replaceOne(filter, newDoc) replaces the entire document (except _id) with newDoc, so any field not included is gone. Passing a document without operators to updateOne is an error in modern drivers, precisely to prevent accidental replacements.

What is an upsert?

An update with upsert: true that inserts a new document if the filter matches nothing. It's how you implement "create or update" in one atomic operation. $setOnInsert sets fields only when the upsert inserts, which is handy for createdAt. A known gotcha: two concurrent upserts with the same filter can both insert unless a unique index on the filter fields prevents it.

How do you paginate, and why is skip() a problem?

skip(n).limit(k) is simple, but the server still walks past n documents, so deep pages get slower linearly. For large datasets or infinite scroll, use range-based (keyset) pagination: remember the last value of an indexed sort key and query for values beyond it.

db.posts
  .find({ _id: { $lt: lastSeenId } })
  .sort({ _id: -1 })
  .limit(20);

Mention the trade-off: range-based pagination can't jump to "page 37" directly. Skip/limit vs. range-based pagination covers both approaches in detail.

Indexing

What types of indexes does MongoDB support?

Single field, compound, multikey (automatically created when you index an array field), text, 2dsphere (geospatial), hashed (mostly for sharding), wildcard (for unpredictable field names), and clustered collections. Index options include unique, partial, sparse, TTL, case-insensitive (via collation), and hidden. You don't need to list all of them; knowing when you'd use unique, compound, partial, and TTL indexes is what matters.

How do you decide the field order in a compound index?

Use the ESR guideline: fields matched by Equality first, then fields used for Sorting, then fields filtered by Range. For the query { status: "open", total: { $gt: 100 } } sorted by createdAt: -1, the index is { status: 1, createdAt: -1, total: 1 }. Putting the range field before the sort field forces an in-memory sort.

Also mention prefixes: an index on { a: 1, b: 1, c: 1 } supports queries on a, on a and b, and on all three, but not on b alone.

What is a covered query?

A query where every field in the filter and the projection is in the index, so MongoDB can answer from the index alone without fetching documents. You typically need to exclude _id from the projection unless it's in the index. In explain() output, a covered query shows totalDocsExamined: 0.

How do you know if a query is using an index?

Run explain("executionStats") and look at the winning plan. IXSCAN means an index was used; COLLSCAN means a full collection scan. Compare totalKeysExamined and totalDocsExamined with nReturned. A well-indexed query examines roughly as many documents as it returns. A SORT stage means the sort happened in memory instead of using index order.

What's the cost of adding indexes?

Every index slows down writes (each insert or update must update every affected index), uses disk space, and competes for RAM in the WiredTiger cache. Unused indexes are pure cost. You can find them with $indexStats, and hide them before dropping to test the impact safely.

Aggregation

Walk me through a basic aggregation pipeline

Documents flow through stages in order, and each stage transforms the stream. A typical pipeline:

db.orders.aggregate([
  { $match: { status: "paid", createdAt: { $gte: ISODate("2026-01-01") } } },
  {
    $group: {
      _id: "$customerId",
      spent: { $sum: "$total" },
      orders: { $sum: 1 },
    },
  },
  { $sort: { spent: -1 } },
  { $limit: 10 },
  {
    $lookup: {
      from: "customers",
      localField: "_id",
      foreignField: "_id",
      as: "customer",
    },
  },
  { $unwind: "$customer" },
  { $project: { _id: 0, name: "$customer.name", spent: 1, orders: 1 } },
]);

Point out that $match comes first so it can use an index and reduce the documents flowing through later stages, and that $lookup comes after $limit so it runs only ten times instead of once per customer.

How does $lookup work, and when should you avoid it?

$lookup performs a left outer join to another collection in the same database, adding matching documents as an array field. It's useful for occasional joins, reports, and admin views. You should avoid depending on it for your hottest read paths: if every page load joins three collections, the data model probably should embed or duplicate some of that data instead. The foreign field should be indexed, or each lookup scans the other collection.

What is $facet used for?

Running several sub-pipelines on the same input documents in one pass, returning each result as a field. The classic use is a search results page that needs the result list, total count, and filter counts (by brand, by price range) in a single query.

What are window functions in MongoDB?

$setWindowFields (MongoDB 5.0+) computes values over a window of related documents without collapsing them like $group does. Examples: running totals, moving averages, and ranks with $rank, $denseRank, and $documentNumber.

db.sales.aggregate([
  {
    $setWindowFields: {
      partitionBy: "$storeId",
      sortBy: { day: 1 },
      output: {
        runningTotal: {
          $sum: "$amount",
          window: { documents: ["unbounded", "current"] },
        },
      },
    },
  },
]);

Data Modeling

When should you embed, and when should you reference?

This is the most important modeling question, and interviewers want reasoning, not a rule. Embed when the data is read together, belongs to one parent, and is bounded in size (a user's addresses, an order's line items). Reference when the related data is large or unbounded, shared across many parents, or updated independently and frequently (a product referenced by many orders, comments on a viral post).

A strong answer mentions the guiding principle, data that is accessed together should be stored together, and hybrid approaches like the subset pattern (embed the latest few items, reference the rest) and extended reference (embed a copy of frequently needed fields, like a customer's name on an order).

What's the problem with unbounded arrays?

Arrays that grow without limit eventually approach the 16 MB document limit, but they cause trouble long before that: every update rewrites a larger document, multikey indexes grow with every element, and reads transfer data you don't need. The fix is to move the growing data into its own collection, bucket it (for example, one document per hour of sensor readings), or keep only a bounded subset embedded.

Name a few MongoDB schema design patterns

Useful ones to know: attribute (key-value arrays for varied, sparse fields, indexed with one compound multikey index), bucket (grouping many small items, like time-series readings, into fewer documents), computed (storing precomputed values such as totals or averages), subset, extended reference, outlier (handling a few unusually large cases differently), and schema versioning (a schemaVersion field to manage migrations).

How do you handle many-to-many relationships?

Usually with arrays of references on one or both sides. For students and courses, each student document might hold an array of course IDs, with a multikey index for "which students take this course?". If both sides are large, a separate "enrollments" collection (like a junction table) with one document per pair and indexes on both keys is cleaner. The right choice depends on cardinality and which direction you query most.

Replication and Sharding

What is a replica set?

A group of mongod processes that hold the same data: one primary that accepts writes, and secondaries that replicate the primary's oplog and can serve reads. If the primary fails, the remaining members hold an election and promote a secondary, usually within seconds. The recommended minimum is three data-bearing members, so a majority survives one failure. Replica sets provide high availability and are also required for transactions and change streams.

What is the oplog?

The operations log: a special capped collection (local.oplog.rs) where the primary records every write as an idempotent operation. Secondaries tail it and apply the same operations. Its size determines the replication window: how long a secondary can fall behind or be offline and still catch up without a full resync.

What are read preference and write concern?

Write concern controls how many members must acknowledge a write before it's considered successful. w: 1 means the primary only; w: "majority" (the default in modern versions) means a majority of voting members, which protects against losing the write in a failover. Read preference controls which members serve reads: primary (default), primaryPreferred, secondary, secondaryPreferred, or nearest. Reading from secondaries can return stale data, since replication is asynchronous.

Bonus: read concern controls the consistency of the data returned (local, majority, snapshot, linearizable).

What is sharding, and what makes a good shard key?

Sharding splits a collection's data across multiple replica sets (shards) based on a shard key, with mongos routers directing queries to the right shards. A good shard key has:

  • High cardinality, so data can be split into many chunks.
  • Even write distribution, avoiding a single "hot" shard. Monotonically increasing keys like timestamps or default ObjectIds send all inserts to one shard, unless you use a hashed shard key.
  • Query isolation, meaning your common queries include the shard key so they target one shard instead of broadcasting to all.

A compound key like { tenantId: 1, createdAt: 1 } often balances these. Since MongoDB 5.0 you can reshard a collection, but it's expensive, so choosing well up front still matters. How to choose a good shard key goes deeper on this.

When would you shard?

When a single replica set can't handle the workload: the dataset or working set exceeds what fits economically on one machine, or write throughput exceeds what one primary can handle. Sharding adds operational complexity, so vertical scaling, better indexes, and archiving old data should come first.

Transactions and Consistency

Does MongoDB support ACID transactions?

Yes. Operations on a single document have always been atomic, including updates to nested fields and arrays. Since 4.0 (replica sets) and 4.2 (sharded clusters), MongoDB supports multi-document ACID transactions. Transactions have a default 60-second lifetime and add overhead, so good schema design that keeps related data in one document means you need them less often. When you do use them, the driver's withTransaction helper handles retries of transient errors.

const session = client.startSession();
await session.withTransaction(async () => {
  await accounts.updateOne(
    { _id: "a" },
    { $inc: { balance: -50 } },
    { session },
  );
  await accounts.updateOne(
    { _id: "b" },
    { $inc: { balance: 50 } },
    { session },
  );
});
await session.endSession();

How do you prevent race conditions without transactions?

Use atomic update operators with conditions in the filter. To decrement stock only if it's available:

db.inventory.updateOne({ _id: sku, qty: { $gte: 1 } }, { $inc: { qty: -1 } });

If modifiedCount is 0, the stock ran out. This check-and-update is atomic on a single document, so no two requests can both take the last unit. For optimistic concurrency on whole documents, include a version field in the filter and increment it on each update.

What are change streams?

A way to subscribe to real-time changes (inserts, updates, deletes) on a collection, database, or deployment, built on the oplog. They're resumable via a resume token, so a consumer that restarts can continue where it left off. Common uses: cache invalidation, syncing to search indexes, event-driven workflows, and audit logs. They require a replica set or sharded cluster.

Scenario Questions

"This query is slow. How do you investigate?"

Walk through a process: reproduce it, run explain("executionStats"), check for COLLSCAN or in-memory SORT, compare documents examined with documents returned, and check whether an index exists and whether its field order follows ESR. Then look wider: the profiler or Atlas Query Insights for patterns, whether the working set fits in RAM, and whether the query returns more data than needed (add a projection). Mentioning that you'd verify the fix with explain again shows rigor.

"Design the schema for a blog with posts, comments, and authors."

A reasonable answer: posts embed a small author summary (name, avatar) as an extended reference, with authorId for the full profile. Comments go in their own collection with postId, indexed on { postId: 1, createdAt: -1 }, because they're unbounded. Posts store a commentCount (computed pattern) and optionally the three latest comments (subset pattern) for list views. Then explain the trade-off: if an author changes their name, you update denormalized copies with an updateMany, which is acceptable because names change rarely.

"Your app gets 'too many connections' errors after moving to serverless. Why?"

Each function instance creates its own MongoClient and connection pool. Under load, hundreds of instances each open several connections. Fixes: create the client once outside the handler so warm invocations reuse it, reduce maxPoolSize, and set sensible idle timeouts. Mention that per-request clients are a bug even outside serverless. (There's a full walkthrough in using MongoDB in serverless functions.)

Tips for the Interview

Talk in trade-offs. "It depends" is only a good answer if you immediately say what it depends on: read/write ratio, data size, consistency needs, and which queries are hot.

Use numbers. 16 MB document limit, 60-second transaction default, three-member replica sets, 100 MB per-stage memory limit for blocking aggregation stages. Concrete limits show real experience.

Admit what's version-specific. Features like window functions, resharding, and queryable encryption arrived in specific releases. Saying "in recent versions" is fine; confidently giving the wrong version isn't.

Write the query. If you're asked how you'd do something, sketch the actual query or pipeline. It's more convincing than a description, and it surfaces details (like $elemMatch) that show you've done the work.

Know what's deprecated. The legacy mongo shell is gone (use mongosh), cursor count() is deprecated in favor of countDocuments() and estimatedDocumentCount(), and some Atlas services have been retired. Up-to-date knowledge signals you work with it now, not five years ago.

Conclusion

MongoDB interviews reward understanding over memorization. Know the fundamentals (documents, BSON, ObjectId), be fluent with queries and array semantics, understand indexes well enough to design one with ESR and verify it with explain(), be comfortable writing aggregation pipelines, and above all be able to reason about embedding versus referencing, shard keys, and consistency trade-offs out loud.

Before your interview, pick three questions from the data modeling and scenario sections, answer them aloud as if to an interviewer, then open mongosh and actually build the schema and queries you described. Anything that doesn't work the way you said it would is exactly what you should study next.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading