
MongoDB Performance Tuning: Memory, WiredTiger Cache, and Disk I/O
You've added the right indexes, your queries use them, and explain() looks clean. Yet at peak traffic, latency still climbs and the server feels sluggish. When that happens, the problem is rarely the query planner. It's usually the layer underneath: how much data fits in memory, how hard the storage engine is working to keep it there, and how fast the disk can serve everything that doesn't fit.
MongoDB's default storage engine, WiredTiger, manages its own cache and relies on the operating system for everything else. Tuning it isn't about flipping a magic setting. It's about understanding where your working set lives, spotting the signs of cache pressure, and making sure the disk and OS aren't sabotaging you. Get this layer right and many "mysterious" slowdowns simply disappear.
This guide covers how MongoDB uses memory, how to size and monitor the WiredTiger cache, how eviction and checkpoints affect latency, and the disk and OS settings that matter most in production.
How MongoDB Uses Memory
A mongod process uses memory in three main ways:
- The WiredTiger internal cache. This holds uncompressed documents and index pages that are actively being read or modified. It's the most important memory region for performance.
- The operating system's filesystem cache. WiredTiger's data files on disk are compressed. When the OS reads them, it keeps those compressed blocks in its page cache. A page that isn't in the WiredTiger cache but is in the filesystem cache can be loaded without touching disk, just decompressed.
- Everything else. Connections (each one costs roughly 1 MB of stack and buffers), aggregation stages that buffer data, in-memory sorts, index builds, and the allocator's own overhead.
That second layer is why MongoDB deliberately doesn't claim all of the RAM. Leaving room for the filesystem cache effectively gives you a second, compressed tier of memory for free.
The Default Cache Size
By default, the WiredTiger cache is the larger of:
- 50% of (total RAM minus 1 GB), or
- 256 MB.
On a 16 GB server that's about 7.5 GB. On a 64 GB server it's about 31.5 GB. In recent versions, mongod also detects container memory limits (cgroups), so a container limited to 4 GB gets a cache of about 1.5 GB rather than half of the host's RAM.
You can check what your server actually configured:
const cache = db.serverStatus().wiredTiger.cache;
const gb = (b) => (b / 1024 ** 3).toFixed(2) + " GB";
printjson({
configured: gb(cache["maximum bytes configured"]),
inCache: gb(cache["bytes currently in the cache"]),
dirty: gb(cache["tracked dirty bytes in the cache"]),
});
{ "configured": "7.50 GB", "inCache": "6.02 GB", "dirty": "0.11 GB" }
Should You Change It?
Usually not. The default is a good balance for a dedicated database server. The cases where changing it makes sense:
- Multiple
mongodprocesses on one host (for example, in a test environment). Each one assumes it owns the machine, so they'll collectively overcommit memory. Divide the cache explicitly. - Other memory-hungry processes on the same host. Not recommended in production, but if you must, shrink the cache to leave room.
- Containers without enforced limits. If the container runtime doesn't expose a cgroup limit,
mongodsees the host's RAM. Set the cache explicitly.
Set it in the config file:
storage:
wiredTiger:
engineConfig:
cacheSizeGB: 4
Resist the urge to crank it up to 80% or 90% of RAM. You'll starve the filesystem cache and connection overhead, and on Linux you risk the OOM killer terminating mongod under load. Bigger is not automatically better.
Understanding Your Working Set
The working set is the portion of your data and indexes that your application touches regularly. Performance is best when the working set fits in the WiredTiger cache. It degrades gracefully when it spills into the filesystem cache, and it falls off a cliff when it has to come from disk on every request.
Indexes are the part you should care about most. If every query needs to walk an index, that index needs to stay hot. Check index sizes across a database with $collStats:
db.getCollectionNames().forEach((name) => {
const [s] = db
.getCollection(name)
.aggregate([{ $collStats: { storageStats: { scale: 1024 * 1024 } } }])
.toArray();
print(
`${name.padEnd(20)} data=${s.storageStats.size}MB ` +
`indexes=${s.storageStats.totalIndexSize}MB`,
);
});
orders data=18432MB indexes=2210MB
customers data=1204MB indexes=388MB
events data=41210MB indexes=6120MB
Note that size here is the uncompressed data size, which is what occupies the cache. If total index size alone approaches your cache size, you've found your problem before looking at anything else.
Keep in mind that the working set is usually smaller than the full dataset. An events collection with 40 GB of history may only have the last few days of data being read. That's why time-based access patterns and indexes that align with them (queries sorted by recent createdAt, for example) are so cache-friendly.
Finding Unused Indexes
Every index consumes cache and slows writes, whether you query it or not. $indexStats shows how often each index has been used since the last restart:
db.orders.aggregate([
{ $indexStats: {} },
{ $project: { name: 1, ops: "$accesses.ops", since: "$accesses.since" } },
{ $sort: { ops: 1 } },
]);
[
{
name: "legacyStatus_1",
ops: Long("0"),
since: ISODate("2026-08-01T04:12:09Z"),
},
{
name: "customerId_1_createdAt_-1",
ops: Long("9120443"),
since: ISODate("2026-08-01T04:12:09Z"),
},
];
An index with zero ops after weeks of normal traffic is a strong candidate for removal. Hide it first with db.orders.hideIndex("legacyStatus_1"), watch for regressions, and drop it only when you're confident. For more on reading query behavior, see Using explain() to Analyze and Debug Slow MongoDB Queries.
Cache Eviction: Where Latency Hides
WiredTiger keeps the cache from overflowing through eviction: removing clean pages and writing out dirty ones. Background eviction threads do this work normally. The key thresholds are:
| Threshold | Default | What happens |
|---|---|---|
| Eviction target | 80% | Background threads start evicting to keep usage here |
| Eviction trigger | 95% | Application threads are drafted to help evict |
| Dirty target | 5% | Background threads start writing dirty pages |
| Dirty trigger | 20% | Application threads are drafted to help write dirty data |
The rows that matter are the triggers. When cache usage hits 95%, or dirty data hits 20%, WiredTiger starts making application threads (the ones serving your queries) do eviction work before they can proceed. That's when p99 latency spikes and throughput stalls, even though CPU looks fine.
A cache that's 80% full is normal and healthy. A cache that keeps bouncing off 95% is a problem.
Reading Eviction Metrics
A quick snapshot shows how the cache is behaving:
const c = db.serverStatus().wiredTiger.cache;
const max = c["maximum bytes configured"];
printjson({
fillPct: ((c["bytes currently in the cache"] / max) * 100).toFixed(1),
dirtyPct: ((c["tracked dirty bytes in the cache"] / max) * 100).toFixed(1),
pagesReadIntoCache: c["pages read into cache"],
pagesEvictedByAppThreads: c["pages evicted by application threads"],
});
pages read into cache is a counter, so sample it twice a few seconds apart. A steadily high rate means the working set doesn't fit and pages are constantly being loaded. A non-trivial and growing pages evicted by application threads means your queries are paying the eviction tax directly.
mongostat gives you a live view of the same thing:
mongostat --uri "mongodb://monitor:secret@db1.internal:27017/?authSource=admin" 5
insert query update delete getmore command dirty used flushes vsize res qrw arw
212 1840 390 *0 120 412|0 3.1% 79.8% 0 11.2G 8.1G 0|0 2|1
198 1792 402 *0 118 398|0 4.6% 88.2% 0 11.2G 8.1G 0|0 3|0
231 1905 377 *0 131 420|0 19.4% 94.7% 1 11.2G 8.1G 12|8 128|40
The last row tells a story: dirty data near 20%, used cache near 95%, and a queue of readers and writers (qrw) building up. That's a server whose cache can't keep up with the write load.
Fixing Cache Pressure
When eviction is hurting you, the fixes fall into a few buckets:
- Shrink the working set. Drop unused indexes, project only the fields you need, archive old data, and avoid queries that scan large ranges of cold documents.
- Reduce write amplification. Updating a single field of a huge document dirties the whole document in cache. Smaller documents and targeted updates (
$seton specific fields, not full replacements) produce less dirty data. - Add memory. Sometimes the honest answer is that the working set grew and the server didn't. Scaling RAM is often cheaper than weeks of optimization.
- Scale out. Sharding splits the working set across multiple servers, each with its own cache.
Tuning eviction thread counts or thresholds via wiredTigerEngineRuntimeConfig is possible, but treat it as a last resort, ideally with guidance from MongoDB support. The defaults are well tested, and the root cause is almost always a working set that's too big.
Checkpoints and the Journal
Every 60 seconds, WiredTiger writes a checkpoint: a consistent snapshot of all data to disk. Between checkpoints, durability comes from the journal, a write-ahead log flushed to disk frequently (and immediately for writes with j: true or the default majority write concern on replica sets).
Checkpoints produce a burst of write I/O. On a well-provisioned disk you won't notice. On an undersized one, you'll see a periodic latency sawtooth that lines up every minute, which is one of the clearest disk-bottleneck signatures you'll find. The flushes column in mongostat shows when checkpoints happen, so correlating it with latency spikes is easy.
The fix is almost never to change the checkpoint interval. It's to give the disk more IOPS or throughput, or to reduce the volume of dirty data being produced.
Disk I/O: Measuring the Bottleneck
When the working set doesn't fit in memory, disk performance becomes your database performance. On Linux, iostat is the fastest way to see what the disk is doing:
iostat -xm 5 /dev/nvme1n1
Device r/s w/s rMB/s wMB/s r_await w_await aqu-sz %util
nvme1n1 2840.0 910.0 44.20 38.10 4.12 9.80 21.3 99.6
What to look for:
%utilnear 100% means the device is constantly busy. On modern NVMe that's less meaningful than it used to be (they handle parallel requests), so combine it with the next two.r_await/w_awaitare average latencies in milliseconds. Local NVMe should be well under 1 ms. Network-attached cloud volumes are commonly 1 to 3 ms. Sustained values above 10 ms mean the disk is saturated.aqu-szis the average queue size. A consistently deep queue means requests are waiting.
Cloud Volumes and Provisioned IOPS
On cloud providers, disk performance is often a function of volume type and size rather than the hardware. A gp3 volume on AWS, for instance, has a baseline IOPS and throughput that you can raise independently. Many "MongoDB is slow" tickets on cloud VMs turn out to be a volume hitting its provisioned ceiling. Check your provider's volume metrics for throttling or burst-balance exhaustion before assuming the database is at fault.
OS and Filesystem Settings That Matter
MongoDB publishes production notes for each platform, and they're worth reading in full. These are the settings that most commonly bite people on Linux.
Use XFS. MongoDB strongly recommends XFS for WiredTiger data volumes. ext4 has shown performance problems with WiredTiger's access patterns under load.
Mount with noatime. Updating access times on every read is pure overhead for data files:
# /etc/fstab
/dev/nvme1n1 /var/lib/mongo xfs defaults,noatime 0 0
Keep readahead small. WiredTiger reads relatively small, random blocks. Large readahead values pull in data you don't need and waste the filesystem cache. The production notes recommend a small value (in the range of 8 to 32 sectors):
sudo blockdev --getra /dev/nvme1n1 # check current value
sudo blockdev --setra 32 /dev/nvme1n1
Persist it with a udev rule or tuned profile, since blockdev settings don't survive a reboot.
Minimize swapping. Set vm.swappiness to 1 (or a similarly low value) so the kernel strongly prefers dropping filesystem cache over swapping out mongod memory:
echo "vm.swappiness = 1" | sudo tee /etc/sysctl.d/90-mongodb.conf
sudo sysctl --system
Check the Transparent Huge Pages guidance for your version. For years, the advice was to disable THP. MongoDB 8.0 moved to an upgraded TCMalloc allocator and changed that recommendation for Linux, so the right setting depends on your server version. Follow the production notes for the exact version you run rather than copying an old tuning script.
Handle NUMA correctly. On multi-socket servers, run mongod with memory interleaving (for example via numactl --interleave=all) and disable zone reclaim (vm.zone_reclaim_mode = 0). The production notes cover the systemd details.
Raise ulimits. mongod needs many open files and processes. The official packages set sensible limits in the systemd unit, but custom installs often don't. db.serverStatus().connections and startup warnings in the log will flag low limits.
Startup warnings are worth checking after any change. MongoDB logs a warning for many of these misconfigurations:
db.adminCommand({ getLog: "startupWarnings" });
Compression Trade-offs
WiredTiger compresses collections with snappy by default and uses prefix compression for indexes. Snappy is fast and gives moderate compression. zstd usually compresses noticeably better at a modest CPU cost, which can help when disk throughput is the bottleneck, because fewer bytes need to move.
You can set the compressor per collection:
db.createCollection("events", {
storageEngine: {
wiredTiger: { configString: "block_compressor=zstd" },
},
});
Or globally for new collections:
storage:
wiredTiger:
collectionConfig:
blockCompressor: zstd
Compression only affects data on disk and in the filesystem cache. Documents in the WiredTiger cache are always uncompressed, so better compression won't make more documents fit in the WiredTiger cache itself, but it does make the filesystem cache and disk go further.
A Practical Tuning Checklist
Start with the working set, not the knobs. Compare total index size and hot data size against the cache. If indexes don't fit, no OS tweak will save you.
Watch the triggers, not the averages. A cache hovering at 80% is fine. Sustained time above 95% fill or 20% dirty is where application threads start doing eviction and latency suffers.
Correlate latency with checkpoints. A spike every 60 seconds points at disk write capacity.
Measure the disk directly. iostat await times and cloud volume throttling metrics tell you whether the hardware is the limit.
Fix the OS basics once. XFS, noatime, low readahead, low swappiness, correct NUMA and THP settings for your version. Bake them into your provisioning so every node gets them.
Don't oversize the cache. Leave headroom for the filesystem cache, connections, and aggregation memory. An OOM-killed primary is far worse than a slightly smaller cache.
Monitor continuously. One-off checks catch today's problem. Trending cache fill, dirty percentage, and disk latency over weeks shows you when growth is about to catch up with your hardware. Monitoring MongoDB with Prometheus and Grafana walks through setting that up.
Conclusion
MongoDB performance at scale is mostly a memory and disk story. The WiredTiger cache holds your hot documents and indexes uncompressed, the filesystem cache holds compressed data behind it, and the disk catches everything else. When the working set outgrows the cache, eviction pressure pushes work onto your query threads, and disk latency becomes query latency.
The good news is that the diagnosis is concrete: compare index and data sizes to the cache, watch fill and dirty percentages against the 95% and 20% triggers, and read disk await times. Most fixes are equally concrete: drop unused indexes, reduce dirty data, fix OS settings, and add RAM or IOPS when growth demands it.
Your next step: run the cache snapshot script from this guide on your busiest node during peak traffic, then run $indexStats on your largest collection. Between the two, you'll know whether you have a memory problem, and which index to hide first.


