Type something to search...
MongoDB for Beginners: Understanding Documents, Collections, and Databases

MongoDB for Beginners: Understanding Documents, Collections, and Databases

If you learned databases through SQL, your mental model probably looks like a spreadsheet: tables with fixed columns, rows that all share the same shape, and foreign keys stitching everything together. That model is powerful, but it also means that a single "customer" in your application is often scattered across five tables and reassembled with joins every time you read it.

MongoDB takes a different approach. Instead of rows, it stores documents: self-contained records that look a lot like the JSON objects your application already works with. A customer, their addresses, and their preferences can live together in one place, shaped the way your code uses them. Documents are grouped into collections, and collections live inside databases.

This guide covers what each of those three building blocks actually is, how they relate to relational concepts you may already know, how to create and inspect them in mongosh, and the naming rules and design habits that will save you trouble later.

The Big Picture: A Three-Level Hierarchy

A MongoDB deployment (a single server, a replica set, or a sharded cluster) holds one or more databases. Each database holds collections. Each collection holds documents.

deployment
└── database: shop
    ├── collection: customers
    │   ├── { _id: ..., name: "Ada Lovelace", ... }
    │   └── { _id: ..., name: "Alan Turing", ... }
    └── collection: orders
        ├── { _id: ..., customerId: ..., total: 42.5 }
        └── ...

If you want a rough translation from relational terms, it looks like this:

Relational (SQL)MongoDB
DatabaseDatabase
TableCollection
RowDocument
ColumnField
Primary key_id field
JoinEmbedding, or $lookup when needed
Schema (DDL)Optional schema validation

The table is useful for orientation, but don't lean on it too hard. The biggest shift isn't the vocabulary. It's that a document can contain nested objects and arrays, so a lot of data that would need separate tables in SQL naturally lives inside a single document in MongoDB.

Documents: The Unit of Data

A document is an ordered set of field and value pairs. In mongosh and most drivers, you write documents using JSON-like syntax:

{
  _id: ObjectId("66f1c2a9e4b0a1b2c3d4e5f6"),
  name: "Ada Lovelace",
  email: "ada@example.com",
  signedUpAt: ISODate("2026-09-01T10:15:00Z"),
  plan: "pro",
  address: {
    street: "12 St James's Square",
    city: "London",
    country: "UK"
  },
  tags: ["early-adopter", "newsletter"],
  loginCount: 37
}

A few things stand out compared to a SQL row:

  • Nested objects. address is a sub-document with its own fields. You can query it directly with dot notation, like "address.city": "London".
  • Arrays. tags holds multiple values in one field. MongoDB can index and query array elements without a separate join table.
  • Rich types. signedUpAt is a real date, not a string. _id is an ObjectId. MongoDB stores documents as BSON, a binary format that supports more types than plain JSON, including dates, 64-bit integers, decimals, and binary data. (If you want the details, see Understanding BSON: How MongoDB Stores Your Data.)

The _id Field

Every document must have an _id field, and its value must be unique within the collection. It acts as the primary key, and MongoDB automatically creates a unique index on it.

If you insert a document without an _id, the driver (or the server) generates an ObjectId for you. An ObjectId is a 12-byte value that includes a timestamp, so documents created later generally have larger ids. You're free to use your own values instead, such as a string SKU or an integer, as long as they're unique:

db.products.insertOne({ _id: "SKU-1042", name: "Desk Lamp", price: 49 });

The one thing you can't do is change a document's _id after it's inserted. If you need a different id, you have to insert a new document and delete the old one.

Document Size Limit

A single document can be at most 16 MB in BSON form. That's a lot of text, but it's a real ceiling, and it matters for design. If a document contains an array that grows forever (every comment ever posted on a popular article, every sensor reading from a device), you'll eventually hit the limit, and long before that, your reads and updates get slower because you're moving a huge document around. Keep unbounded lists in their own collection.

Collections: Groups of Related Documents

A collection is a group of documents, usually representing one kind of thing: users, orders, products, events. It plays roughly the role of a table, with one important difference: documents in the same collection don't have to share the same fields.

db.products.insertMany([
  { name: "Desk Lamp", price: 49, wattage: 8 },
  { name: "Notebook", price: 6, pages: 120, ruled: true },
  { name: "Gift Card", price: 25, digital: true },
]);

All three are products, but each carries the fields that make sense for it. In SQL you'd either add nullable columns for every possibility or split the data into subtype tables. In MongoDB you just store what you have.

This flexibility is sometimes described as "schemaless", which is misleading. Your application still has a schema; it's simply enforced in code instead of the database by default. And when you want the database to enforce rules, you can add schema validation with JSON Schema to reject documents that don't match. Flexible by default, strict when you choose.

Collections Are Created Lazily

You don't need to create a collection before using it. The first time you insert a document into a collection that doesn't exist, MongoDB creates it:

use shop
db.reviews.insertOne({ productId: "SKU-1042", rating: 5, text: "Great lamp." })
{
  acknowledged: true,
  insertedId: ObjectId('66f1d0b7c1a2b3c4d5e6f701')
}

The reviews collection now exists. You can also create a collection explicitly, which is useful when you want options like validation rules, a capped size, or a time series configuration:

db.createCollection("auditLog", {
  validator: {
    $jsonSchema: {
      bsonType: "object",
      required: ["action", "at"],
      properties: {
        action: { bsonType: "string" },
        at: { bsonType: "date" },
      },
    },
  },
});

Special Kinds of Collections

Most of the time you'll use regular collections, but it helps to know the others exist:

  • Time series collections store measurements over time (metrics, IoT readings) in a compressed, bucketed format.
  • Capped collections have a fixed size and overwrite their oldest documents when full.
  • Views are read-only collections defined by an aggregation pipeline over another collection.

You don't need any of these on day one. Just know that "collection" covers more than a plain bag of documents.

Databases: Namespaces for Collections

A database is a container for collections. It's mostly a namespace: it groups related collections together, and it's the level at which you commonly assign user permissions. One MongoDB deployment can host many databases, and a typical application uses one database per app or per environment.

show dbs
admin    40.00 KiB
config  108.00 KiB
local    72.00 KiB
shop     96.00 KiB

Like collections, databases are created lazily. Running use blog switches your shell context to a database named blog, but nothing is written to disk until you insert data into one of its collections. That's why a brand new database doesn't appear in show dbs until it has something in it.

Reserved Databases

Three databases have special meaning, and you shouldn't store application data in them:

  • admin holds users and roles with cluster-wide privileges and is used for administrative commands.
  • local stores data specific to a single server, including the replication oplog. It's never replicated.
  • config is used internally, especially by sharded clusters, to store metadata.

Namespaces

The combination of database and collection name, like shop.orders, is called a namespace. You'll see this term in logs, error messages, and tools like the profiler. When an error says something about shop.orders, it's telling you exactly which collection was involved.

Putting It Together in mongosh

Let's walk through a small session end to end. Start the shell against a local server (or an Atlas connection string):

mongosh "mongodb://localhost:27017"

Switch to a database and insert a few documents:

use library

db.books.insertMany([
  {
    title: "The Pragmatic Programmer",
    authors: ["David Thomas", "Andrew Hunt"],
    year: 1999,
    genres: ["software", "career"],
    available: true
  },
  {
    title: "Designing Data-Intensive Applications",
    authors: ["Martin Kleppmann"],
    year: 2017,
    genres: ["software", "databases"],
    available: false
  }
])

List what now exists:

show collections
books

Query using the document structure directly. Matching a value inside an array works without any special syntax:

db.books.find({ genres: "databases" }, { title: 1, year: 1, _id: 0 });
[ { title: 'Designing Data-Intensive Applications', year: 2017 } ]

And count documents with a filter:

db.books.countDocuments({ available: true });
1

Everything you just did (creating the database, creating the collection, inserting, querying) took no upfront schema definition. That's the day-to-day feel of MongoDB. For a full tour of inserting, reading, updating, and deleting, see MongoDB CRUD Operations Explained with Practical Examples.

Naming Rules and Conventions

MongoDB is permissive about names, but a few rules are enforced and a few conventions will keep you out of trouble.

Database names can't contain spaces or characters like / \ . " $, and on some platforms they're case-insensitive on disk, so don't create Shop and shop as separate databases. Keep them short, lowercase, and descriptive.

Collection names can't be empty, can't start with system. (reserved for internal use), and shouldn't contain $. Most teams use lowercase plural nouns (orders, line_items) or camelCase (lineItems). Pick one style and stick to it.

Field names can't be empty, and while recent MongoDB versions technically allow field names containing . and $, they cause endless problems with query syntax and tooling. Avoid them. Also remember that field names are stored in every single document, so a field called customerShippingAddressPostalCode repeated across ten million documents costs real storage. Clear names are still worth it, but don't go overboard.

How to Think About Document Design

The hardest part of learning MongoDB isn't the syntax. It's deciding what goes into a document. A useful guiding principle: data that is accessed together should be stored together.

Consider a blog post with comments. You could model it relationally:

// posts
{ _id: 1, title: "Hello MongoDB", body: "..." }

// comments
{ _id: 101, postId: 1, author: "sam", text: "Nice intro" }
{ _id: 102, postId: 1, author: "lee", text: "Thanks!" }

Or you could embed:

{
  _id: 1,
  title: "Hello MongoDB",
  body: "...",
  comments: [
    { author: "sam", text: "Nice intro" },
    { author: "lee", text: "Thanks!" }
  ]
}

Embedding means one read returns the whole post with its comments, and updates to a single document are atomic. Referencing means comments can grow without bound and be queried independently. Neither is universally right:

  • Embed when the child data is small, bounded, and almost always read with the parent (an order's line items, a user's addresses).
  • Reference when the child data grows without limit, is shared by many parents, or is frequently accessed on its own (comments on a viral post, products referenced by many orders).

A common middle ground is to embed a small summary (say, the three most recent comments) and keep the full list in a separate collection. You'll run into this pattern constantly once you start looking for it.

Common Mistakes

Treating MongoDB like a SQL database with different syntax. If every collection is a normalized table and every read needs three $lookup stages, you're paying for flexibility without using it. Start from how your application reads data, then shape documents to match.

Letting arrays grow forever. Unbounded arrays push documents toward the 16 MB limit and make every update rewrite more data. If a list can grow without a natural cap, give it its own collection.

Storing dates and numbers as strings. "2026-09-15" sorts like text and can't use date operators. "49.99" can't be summed. Use real BSON types from the start; converting millions of documents later is painful.

Inconsistent field names across documents. Flexible doesn't mean chaotic. If half your documents use userId and the other half use user_id, every query has to handle both. Agree on names, and consider schema validation once your model settles.

Creating a database per user or a collection per day. Thousands of databases or collections add overhead for the server (each collection and index has files and metadata). Use a field like tenantId or a date field inside a shared collection instead, and index it.

Conclusion

MongoDB's data model comes down to three layers. Documents are rich, self-contained records with nested objects, arrays, and typed values, each identified by a unique _id. Collections group documents of the same kind without forcing them into identical shapes. Databases group collections into namespaces and permission boundaries. Once you're comfortable with that hierarchy, and with the idea that data accessed together belongs together, the rest of MongoDB starts to make sense.

For a next step, pick one entity from a project you know well, such as a user profile or an order, and sketch it as a single MongoDB document. Decide which related data you'd embed and which you'd reference, then insert it into a local mongosh session and try querying it by a nested field.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading