Type something to search...
Understanding the WiredTiger Storage Engine in MongoDB

Understanding the WiredTiger Storage Engine in MongoDB

When you call insertOne() or run a query, you're talking to MongoDB's query layer: it parses the command, picks an index, and builds a plan. But the part that actually stores bytes on disk, keeps them consistent after a crash, and lets thousands of operations run at once without corrupting each other is a separate component: the storage engine.

Since MongoDB 3.2, the default storage engine has been WiredTiger, and since 4.2 the old MMAPv1 engine is gone entirely. WiredTiger is responsible for document-level concurrency, compression, crash recovery, and much of the behavior people think of as "just how MongoDB works." Understanding it explains why some write patterns are fast and others aren't, why deleted data doesn't shrink your disk usage, and why memory settings matter so much.

This guide covers WiredTiger's architecture: how data is organized on disk, how the cache works, how concurrency and snapshots are handled, what checkpoints and the journal do, how compression is configured, and the practical consequences for developers and operators.

Where WiredTiger Sits in MongoDB

MongoDB has a pluggable storage engine API. The layers look roughly like this:

  1. Query and command layer: parsing, query planning, aggregation, replication logic.
  2. Storage engine API: a generic interface for "store this record," "read this key," "start a transaction."
  3. WiredTiger: the implementation of that interface, managing tables, the cache, transactions, and files.

Every collection and every index in MongoDB is a separate WiredTiger table, stored as its own file. A collection maps record IDs to BSON documents. Each index maps index keys to record IDs. When you query by an indexed field, MongoDB looks up the key in the index table, gets a record ID, and fetches the document from the collection table.

WiredTiger itself knows nothing about BSON, queries, or replica sets. It stores keys and values reliably and concurrently. That separation is why the same engine can serve regular collections, indexes, the oplog, and internal metadata.

What's in the Data Directory

If you look inside a dbPath, you'll see WiredTiger's files:

ls -1 /var/lib/mongo
WiredTiger
WiredTiger.lock
WiredTiger.turtle
WiredTiger.wt
WiredTigerHS.wt
_mdb_catalog.wt
collection-0-4421807362164239184.wt
collection-2-4421807362164239184.wt
diagnostic.data
index-1-4421807362164239184.wt
index-3-4421807362164239184.wt
journal
mongod.lock
sizeStorer.wt

The important ones:

  • collection-*.wt and index-*.wt: one file per collection and per index.
  • _mdb_catalog.wt: MongoDB's catalog, mapping namespaces like shop.orders to WiredTiger table names.
  • WiredTiger.wt and WiredTiger.turtle: WiredTiger's own metadata. The turtle file bootstraps the metadata table.
  • WiredTigerHS.wt: the history store, holding older versions of data needed by long-running snapshots.
  • sizeStorer.wt: cached document counts and data sizes per collection, which is what makes estimatedDocumentCount() instant.
  • journal/: the write-ahead log files.
  • diagnostic.data/: FTDC metrics, captured continuously for support diagnostics.

You can map a collection to its file with $collStats:

db.orders.aggregate([{ $collStats: { storageStats: {} } }]).toArray()[0]
  .storageStats.wiredTiger.uri;
statistics:table:collection-2-4421807362164239184

Never modify, copy, or delete individual .wt files by hand. They only make sense together with the catalog and metadata, and copying a single file into another server won't give you a usable collection.

B-Trees and Pages

WiredTiger stores each table as a B-tree made of pages. On disk, pages are compressed blocks. In memory, they're uncompressed and have a different, update-friendly structure.

Reads walk from the root page down through internal pages to a leaf page holding the key. Because indexes are B-trees too, range scans on an index (like createdAt between two dates) are efficient: once the first key is found, subsequent keys are next to it in the leaf pages.

Writes don't rewrite pages in place. When a document is updated, WiredTiger attaches the new version to the in-memory page as an update record. Later, when the page is reconciled (during eviction or a checkpoint), the updates are merged into a new page image and written to a new location in the file. This copy-on-write approach is what lets checkpoints produce consistent snapshots without blocking writers.

One practical consequence: updating a large document creates a new full version of that document in memory. There's no in-place byte patching. Keep documents reasonably sized and prefer targeted updates, and you'll produce less dirty data for WiredTiger to write out.

The WiredTiger Cache

WiredTiger keeps its own internal cache of uncompressed pages. By default it's sized at the larger of 50% of (RAM minus 1 GB) or 256 MB. Everything else in memory is left to the operating system, which caches the compressed data files.

When a query needs a page that isn't in the cache, WiredTiger reads the compressed block (hopefully from the OS page cache, otherwise from disk), decompresses it, and loads it. When the cache gets full, eviction removes clean pages and writes dirty ones.

Background threads normally handle eviction. When the cache fills past certain thresholds, WiredTiger pulls application threads into eviction work, and that's when latency spikes. Cache sizing, eviction metrics, and disk tuning get a full treatment in MongoDB Performance Tuning: Memory, WiredTiger Cache, and Disk I/O.

Concurrency: Document-Level, Optimistic

One of WiredTiger's biggest wins over MMAPv1 was concurrency. MMAPv1 locked at the collection level for writes. WiredTiger uses document-level concurrency control, so two clients updating different documents in the same collection don't block each other.

It does this optimistically. When two operations try to modify the same document at the same moment, one succeeds and the other gets an internal write conflict. MongoDB transparently retries the conflicting operation for single-document writes, so your application doesn't see an error. Inside multi-document transactions, however, a write conflict aborts the transaction with a WriteConflict error labeled TransientTransactionError, and the driver's withTransaction() helper retries it.

You can see how often this is happening:

db.serverStatus().metrics.operation.writeConflicts;

A steady stream of write conflicts usually points at a hot document: a single counter, a shared status record, or an "inventory" document that every request updates. The fix is a data-modeling change, not a configuration one. Spread the writes (bucketed counters, per-shard counters) or restructure so different requests touch different documents.

Concurrency Tickets

WiredTiger also limits how many operations can be active inside the storage engine at once, using read and write tickets. When all tickets are in use, new operations queue. Historically the limit was a fixed 128 each. In MongoDB 7.0 and later, the server adjusts the number of concurrent storage transactions dynamically based on observed throughput.

You can check ticket usage in serverStatus:

db.serverStatus().queues.execution;

The exact shape of this section has changed across versions, so inspect it on your server rather than hard-coding field names in scripts. If operations are frequently waiting for tickets, the storage engine is saturated, typically because of cache pressure or slow disk, and adding more concurrency won't help.

Snapshots and MVCC

WiredTiger uses multi-version concurrency control (MVCC). Instead of locking a document while someone reads it, it keeps multiple versions of recently changed data. Each operation reads from a snapshot: a consistent point-in-time view of the data.

That's how a long-running query can scan a collection while writes continue: it sees the data as of its snapshot, and writers create new versions rather than overwriting the ones it's reading.

This is also the foundation for:

  • Multi-document transactions. A transaction reads from one snapshot and commits its writes atomically.
  • Read concern "snapshot" and majority reads. MongoDB asks WiredTiger for a snapshot at a specific timestamp, such as the majority-committed point of a replica set.
  • Causal consistency and cluster-time reads.

Old versions can't be discarded while any snapshot still needs them. WiredTiger moves them into the history store (WiredTigerHS.wt), which is why very long-running transactions or cursors with old snapshots can bloat cache and disk usage. MongoDB limits transaction lifetime by default (transactionLifetimeLimitSeconds, 60 seconds) partly for this reason.

Checkpoints

WiredTiger writes a checkpoint every 60 seconds by default. A checkpoint is a consistent on-disk snapshot of all tables. It flushes dirty pages, writes new page images to free locations in the files, and finally updates metadata to point at the new checkpoint.

Because pages are written copy-on-write, the previous checkpoint stays intact until the new one is complete. If the server crashes mid-checkpoint, the old one is still valid. That makes checkpoints a crash-safe baseline: on restart, WiredTiger starts from the last completed checkpoint.

Checkpoints are also a source of periodic I/O bursts. On an undersized disk, you'll see latency rise in a rhythm that matches the checkpoint interval.

The Journal

Checkpoints alone would mean losing up to 60 seconds of writes in a crash. The journal closes that gap. It's a write-ahead log: before (or as) a change is applied in memory, a record of it is appended to the journal files. On restart, WiredTiger loads the last checkpoint and replays the journal to recover everything written after it.

Journal records are batched and flushed to disk frequently. The group-commit interval is controlled by storage.journal.commitIntervalMs (default 100 ms). Writes that request journaling, via j: true or a majority write concern on replica sets (where journaling is implied by default), wait for the journal flush before they're acknowledged:

db.payments.insertOne(
  { orderId: 8812, amount: NumberDecimal("149.00"), status: "captured" },
  { writeConcern: { w: "majority", j: true } },
);

In older versions you could disable journaling. That's no longer possible for replica set members in modern releases; the journal is always on, and that's the correct default. Journal files are compressed with snappy by default, configurable via storage.wiredTiger.engineConfig.journalCompressor.

Compression

WiredTiger compresses data on disk, and the savings are often substantial for JSON-like data with repeated field names.

TargetDefaultAlternatives
Collectionssnappyzstd, zlib, none
Indexesprefix compressioncan be disabled
Journalsnappyzstd, zlib, none

Snappy is fast with moderate compression. zstd typically compresses noticeably better with reasonable CPU cost, and it's a popular choice for large or append-heavy collections like logs and events. zlib compresses well but is slower. Prefix compression for indexes stores shared key prefixes once per page, which works especially well for compound indexes whose keys share a leading field.

Set compression per collection at creation time:

db.createCollection("audit_events", {
  storageEngine: {
    wiredTiger: { configString: "block_compressor=zstd" },
  },
});

Or set the defaults for new collections in mongod.conf:

storage:
  wiredTiger:
    engineConfig:
      journalCompressor: snappy
    collectionConfig:
      blockCompressor: zstd
    indexConfig:
      prefixCompression: true

Changing the config doesn't recompress existing collections. To migrate an existing collection to a different compressor, you need to rewrite its data, for example by creating a new collection with the desired settings and copying documents over, or by resyncing a replica set member after changing the default (initial sync recreates collections with the new setting).

To see how well compression is working, compare the uncompressed data size with storage size:

const { storageStats: s } = db.audit_events
  .aggregate([{ $collStats: { storageStats: {} } }])
  .toArray()[0];

print(`ratio: ${(s.size / s.storageSize).toFixed(2)}x`);
ratio: 5.84x

Why Disk Space Doesn't Shrink

A common surprise: you delete half the documents in a collection, but the file on disk stays the same size. That's by design. WiredTiger marks the freed blocks as available and reuses them for future writes to the same collection. It doesn't return them to the operating system automatically.

You can see how much reusable space a collection has:

const { storageStats: s } = db.events
  .aggregate([{ $collStats: { storageStats: {} } }])
  .toArray()[0];

printjson({
  storageSizeMB: (s.storageSize / 1024 ** 2).toFixed(0),
  reusableMB: (
    s.wiredTiger["block-manager"]["file bytes available for reuse"] /
    1024 ** 2
  ).toFixed(0),
});

If a large share of the file is reusable and you won't refill it soon, the compact command can release space:

db.runCommand({ compact: "events" });

compact is much less disruptive than it used to be in older versions, but it still generates significant I/O. Run it on secondaries first, one member at a time, and step down the primary before compacting it. Recent versions (8.0 and later) also offer an autoCompact command that runs background compaction across the node; check the documentation for your version before enabling it. On a replica set, resyncing a member from scratch is another way to get a compact copy of the data.

For data you routinely delete by age, TTL indexes and time-series collections help keep growth predictable, since freed space is continuously reused by new data.

Other Storage Engines

WiredTiger is the only storage engine in MongoDB Community Edition. MongoDB Enterprise also includes an in-memory storage engine, based on WiredTiger, that keeps all data in memory without persisting it to disk. It's aimed at workloads that need predictable low latency on data that can be rebuilt, and it's rarely the right choice for general application data.

You can confirm which engine a server is running:

db.serverStatus().storageEngine;
{
  name: 'wiredTiger',
  supportsCommittedReads: true,
  persistent: true,
  ...
}

Practical Takeaways for Developers

Avoid hot documents. Document-level concurrency only helps if writes are spread across documents. A single shared counter becomes a write-conflict bottleneck.

Keep documents a sensible size. Updates create new versions of the whole document in the cache. Unbounded arrays and huge embedded blobs amplify every write.

Keep transactions short. Long transactions pin old snapshots, push data into the history store, and increase cache pressure for everyone.

Choose compression deliberately. zstd for large, append-heavy collections can cut disk usage and I/O. Set it at creation time, because changing it later means rewriting data.

Expect disk usage to plateau, not shrink. Deletes free space for reuse. Plan for compact or member resyncs only if you truly need space returned to the OS.

Trust the journal. Use a majority write concern for important data and let the journal provide crash durability. Don't try to work around it for speed.

Conclusion

WiredTiger is the part of MongoDB that turns documents into durable, concurrent, compressed storage. Collections and indexes are separate B-tree tables, an internal cache holds uncompressed pages while the OS caches compressed files, MVCC snapshots let readers and writers coexist, checkpoints provide a consistent on-disk baseline every minute, and the journal covers the gap between them. Compression settings and deletion behavior follow directly from how those pieces work.

Once you understand those mechanics, a lot of operational behavior stops being surprising. Your next step: run $collStats on your three largest collections and compare size to storageSize and to the reusable-bytes figure. You'll learn how well your data compresses, and whether a compressor change or a compaction is worth planning.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading