Type something to search...
One-to-Many Relationships in MongoDB: Patterns and Trade-Offs

One-to-Many Relationships in MongoDB: Patterns and Trade-Offs

In a relational database, a one-to-many relationship has one answer: a foreign key on the "many" side, and a join when you need both. You normalize first and optimize later. MongoDB gives you more options, and that freedom is both its biggest strength and the source of most schema regret. Embed too much and documents bloat past usefulness. Reference everything and you've rebuilt a relational database without joins that are as cheap.

The good news is that the choice isn't a matter of taste. It follows from a few concrete questions: how many children a parent can have, whether children are read with their parent, whether they're shared, and how often they change. Answer those, and the right pattern is usually obvious.

This guide covers the main one-to-many patterns (embedding, child references, parent references), hybrid approaches like the subset and extended reference patterns, how each one handles reads, writes, and growth, and a decision framework you can apply to your own data.

The Questions That Decide the Pattern

Before looking at patterns, get concrete about the relationship. For each one-to-many in your app, ask:

  1. Cardinality. Is "many" a handful, hundreds, or potentially millions? Is there an upper bound you can state with confidence?
  2. Access pattern. When you read the parent, do you almost always need the children too? Do you ever read children without the parent?
  3. Independence. Do children have a life of their own? Are they queried, updated, or listed across parents?
  4. Volatility. How often are children added or changed, and by how many writers at once?

MongoDB's modelling guidance has a useful shorthand for cardinality: one-to-few, one-to-many, and one-to-squillions. They map neatly onto the three core patterns.

Pattern 1: Embedding (One-to-Few)

Store the children as an array inside the parent:

db.users.insertOne({
  _id: 1,
  name: "Ada",
  addresses: [
    { label: "home", street: "12 Analytical Way", city: "London" },
    { label: "work", street: "1 Engine St", city: "Cambridge" },
  ],
});

One read gets everything. One write updates the parent and its children atomically, with no transaction needed. Queries on children are easy with dot notation and $elemMatch:

db.users.find({ "addresses.city": "Cambridge" });

Embedding is the right default when:

  • The number of children is small and bounded (addresses, phone numbers, a recipe's ingredients, an order's line items).
  • Children are almost always read with the parent.
  • Children don't make sense on their own and aren't shared between parents.

Where Embedding Breaks Down

Embedding struggles when the array can grow without limit. Every document has a 16 MB BSON limit, but you'll hit practical problems long before that. Large documents mean more data read from disk and sent over the network for every query, even when you only need the parent's name. Multikey indexes on the array grow with every element. Concurrent $push operations on a single hot document contend with each other.

Order line items embed well because an order has a natural limit and is effectively immutable after purchase. Blog post comments are a classic trap: they seem few at first, but a popular post can collect thousands, and you usually show only the latest dozen.

Pattern 2: Child References (One-to-Many)

Store the children in their own collection and keep an array of their IDs in the parent:

db.parts.insertMany([
  { _id: "P-100", name: "Bolt M6", price: 0.12 },
  { _id: "P-101", name: "Washer M6", price: 0.03 },
  { _id: "P-230", name: "Hinge", price: 2.4 },
]);

db.products.insertOne({
  _id: "CAB-01",
  name: "Wall cabinet",
  parts: ["P-100", "P-101", "P-230"],
});

To read a product with its parts, fetch the product and then the parts:

const product = db.products.findOne({ _id: "CAB-01" });
const parts = db.parts.find({ _id: { $in: product.parts } }).toArray();

Or do it in one pipeline with $lookup:

db.products.aggregate([
  { $match: { _id: "CAB-01" } },
  {
    $lookup: {
      from: "parts",
      localField: "parts",
      foreignField: "_id",
      as: "partDocs",
    },
  },
]);

Child references fit when children are shared between parents (the same bolt appears in many products), when they need to be queried or updated independently, and when the number of children is moderate: dozens to low thousands. The ID array itself is small, so the parent document stays light.

The downside is that keeping the array in sync is your responsibility. Deleting a part doesn't remove it from every product's parts array. And the ID array still grows with the relationship, so it doesn't suit truly unbounded cases.

Pattern 3: Parent References (One-to-Squillions)

Store the parent's ID on each child, and nothing on the parent:

db.posts.insertOne({
  _id: ObjectId("66e7f0a1c2d3e4f5a6b7c8d9"),
  title: "Hello MongoDB",
});

db.comments.insertMany([
  {
    postId: ObjectId("66e7f0a1c2d3e4f5a6b7c8d9"),
    author: "grace",
    text: "Great intro!",
    createdAt: new Date("2026-09-16T09:00:00Z"),
  },
  {
    postId: ObjectId("66e7f0a1c2d3e4f5a6b7c8d9"),
    author: "linus",
    text: "Could you cover indexing next?",
    createdAt: new Date("2026-09-16T09:05:00Z"),
  },
]);

db.comments.createIndex({ postId: 1, createdAt: -1 });

This is the relational model, and it's the right choice when the "many" side is unbounded: comments, log entries, sensor readings, messages, orders per customer. Each child is a small, independent document. Adding one is a simple insert that doesn't touch the parent, so there's no contention and no size ceiling.

Reading is a query on the child collection, and the compound index makes "latest comments for this post" fast:

db.comments
  .find({ postId: ObjectId("66e7f0a1c2d3e4f5a6b7c8d9") })
  .sort({ createdAt: -1 })
  .limit(20);

The index is not optional. Without it, every page view scans the whole comments collection. With it, the query reads exactly the 20 entries it needs.

The trade-off: getting a parent and its children takes two queries (or a $lookup), and there's no single-document atomicity across them. If you need to create a post and its first comment together, all-or-nothing, you need a multi-document transaction.

Comparing the Core Patterns

QuestionEmbedChild refsParent refs
Typical size of "many"Few, boundedDozens to thousandsUnbounded
Read parent + childrenOne readTwo reads or $lookupTwo reads or $lookup
Children shared across parentsNoYesRarely
Query children independentlyAwkwardEasyEasy
Adding a childUpdate parentInsert + update parentInsert only
Atomic parent + child writeYesNeeds transactionNeeds transaction
Growth riskDocument sizeID array sizeNone

Hybrid Patterns

Real apps rarely fit one pattern perfectly. These hybrids are where MongoDB modelling gets practical.

The Subset Pattern

Keep the full list in its own collection (parent references), but embed a small, frequently needed subset in the parent. For a product page that shows the three most recent reviews:

db.products.updateOne(
  { _id: "LAMP-04" },
  {
    $push: {
      recentReviews: {
        $each: [
          {
            user: "ada",
            rating: 5,
            text: "Lovely warm light.",
            at: new Date(),
          },
        ],
        $sort: { at: -1 },
        $slice: 3,
      },
    },
    $inc: { reviewCount: 1, ratingTotal: 5 },
  },
);

db.reviews.insertOne({
  productId: "LAMP-04",
  user: "ada",
  rating: 5,
  text: "Lovely warm light.",
  at: new Date(),
});

The product page renders from one document. The "all reviews" page queries the reviews collection with pagination. The embedded array is capped by $slice, so it can't grow unbounded. The cost is writing twice, which you can wrap in a transaction if the two must never disagree, or accept brief inconsistency if they can.

The Extended Reference Pattern

A reference stores only an ID, which means showing anything about the related document requires a lookup. The extended reference copies the few fields you display alongside the ID:

db.orders.insertOne({
  _id: "ORD-7781",
  customer: {
    _id: ObjectId("66e7f1b2c3d4e5f6a7b8c9d0"),
    name: "Grace Hopper",
    email: "grace@example.com",
  },
  items: [{ sku: "MUG-01", qty: 2, price: 12 }],
  total: 24,
});

The order list can show the customer's name without touching the customers collection. This is deliberate denormalization, and it raises the question of what happens when the source changes. Sometimes the answer is "nothing", and that's correct: an order should record the shipping name at the time of purchase. When the copy must stay current, propagate changes:

db.orders.updateMany(
  { "customer._id": ObjectId("66e7f1b2c3d4e5f6a7b8c9d0") },
  { $set: { "customer.name": "Grace B. Hopper" } },
);

The rule: copy fields that are read often and change rarely. Copying a field that changes every minute creates a write amplification problem.

The Bucket Pattern

For high-volume children like sensor readings or chat messages, storing one document per event creates a huge number of tiny documents. Grouping them into buckets (for example, one document per device per hour) reduces document count and index size:

db.readings.updateOne(
  {
    deviceId: "th-22",
    hour: new Date("2026-09-16T08:00:00Z"),
    count: { $lt: 200 },
  },
  {
    $push: { samples: { t: new Date(), tempC: 21.4 } },
    $inc: { count: 1 },
  },
  { upsert: true },
);

The count condition caps each bucket, and the upsert opens a new one when the current bucket is full. For time-series data specifically, MongoDB's native time series collections do this bucketing for you and are usually the better choice; see the guide to time series collections.

Keeping References Consistent

MongoDB doesn't enforce foreign keys, so referential integrity is the application's job. A few practical tools:

  • Transactions when a parent and child must be created or deleted together.
  • Cascading deletes in code: when deleting a post, delete its comments in the same transaction, or mark them for background cleanup.
  • Schema validation to require that postId exists and is an ObjectId, which catches type mismatches (see the post on schema validation).
  • Consistent types: if posts._id is an ObjectId, comments.postId must be too. A string reference will never match in queries or $lookup.

Here's a transactional delete with the Node.js driver:

const session = client.startSession();
try {
  await session.withTransaction(async () => {
    await db.collection("comments").deleteMany({ postId }, { session });
    await db.collection("posts").deleteOne({ _id: postId }, { session });
  });
} finally {
  await session.endSession();
}

For posts with tens of thousands of comments, a single transaction deleting them all may be slow or hit limits. In that case, soft-delete the post first (so it disappears from the app) and remove comments in batches with a background job.

A Worked Example: A Blog Platform

Applying the questions to a blog with posts, authors, tags, and comments:

  • Post to tags: few, bounded, read with the post. Embed as an array of strings.
  • Post to author: many posts per author, and the post list shows the author's name and avatar. Extended reference on the post: author: { _id, name, avatarUrl }.
  • Post to comments: unbounded, paginated, written concurrently by readers. Parent reference in a comments collection indexed on { postId: 1, createdAt: -1 }, plus a commentCount on the post updated with $inc.
  • Post to revisions: potentially many, rarely read. Parent reference in a revisions collection.

The post document stays small and renders the whole article page with one read, while the high-volume data lives where it can grow.

Common Pitfalls

Embedding unbounded data. If you can't state a maximum for an embedded array, it will eventually cause trouble. Move it to its own collection or cap it with the subset pattern.

Normalizing everything by habit. Splitting an order's line items into a separate collection turns every order view into two queries for no benefit. If data is always read together and bounded, embed it.

Mismatched reference types. Storing customerId as a string while customers._id is an ObjectId silently breaks every lookup. Keep reference types identical to the keys they point to.

Missing indexes on parent references. A postId field without an index means every "comments for this post" query is a collection scan.

Denormalizing volatile fields. Copying fields that change often into thousands of child documents turns every change into a large updateMany. Copy stable fields, reference volatile ones.

Conclusion

One-to-many in MongoDB comes down to three core patterns. Embed when the children are few, bounded, and read with the parent. Use child references when children are shared or independently managed but still moderate in number. Use parent references when the "many" side is unbounded. Then reach for hybrids like the subset and extended reference patterns to make your most common reads a single document fetch, and accept the small cost of keeping copies in sync.

List every one-to-many relationship in your current app, and write down the maximum realistic size of the "many" side next to each. Any embedded array without a confident maximum is the first candidate to move into its own collection.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading