Type something to search...
Finding Slow Queries with the MongoDB Database Profiler

Finding Slow Queries with the MongoDB Database Profiler

Most performance problems don't announce themselves with a single terrible query. They show up as a vague complaint: the dashboard feels sluggish, p99 latency crept up after the last release, the database CPU sits at 70% for no obvious reason. Somewhere in the thousands of operations hitting your cluster every second, a handful are doing far more work than they should. The question is which ones.

The database profiler answers that question. It records details about operations (the full command, how many index keys and documents it examined, which plan it used, how long it took, and which app sent it) into a queryable collection. Combined with the slow query log that every mongod writes, it gives you a ranked list of what to fix, instead of a hunch.

This guide covers profiling levels and thresholds, reading system.profile documents, queries that surface the worst offenders, sizing the profile collection, filtering what gets captured, using the slow query log, how profiling works on replica sets, sharded clusters, and Atlas, and how to keep the overhead under control.

Profiling Levels

The profiler is configured per database and has three levels:

LevelWhat gets recorded
0Nothing in system.profile (the default). Slow operations still go to the log.
1Operations slower than slowms (default 100 ms), optionally sampled or filtered
2Every operation on the database

Level 1 is the one you'll use. Level 2 records everything, which is useful for a few minutes of debugging on a development machine and dangerous on a busy production server.

Check the current settings:

use shop
db.getProfilingStatus()
{ was: 0, slowms: 100, sampleRate: 1, ok: 1 }

Enable level 1 with a custom threshold:

db.setProfilingLevel(1, { slowms: 50 });
{ was: 0, slowms: 100, sampleRate: 1, ok: 1 }

The returned document shows the previous settings, which confuses people the first time. Run db.getProfilingStatus() again to confirm the new ones.

A few details worth knowing:

  • Settings apply per database, except slowms and sampleRate, which are shared across the whole mongod instance. Changing slowms in one database changes it for the log threshold everywhere.
  • Settings don't survive a restart when set with setProfilingLevel. To make them permanent, put them in the config file (shown later).
  • Level 0 still logs slow operations. The slowms threshold controls what goes to the server log even when the profiler itself is off.

Sampling

On a busy server, even level 1 can record a lot. sampleRate records only a fraction of slow operations:

db.setProfilingLevel(1, { slowms: 50, sampleRate: 0.25 });

Now roughly one in four operations slower than 50 ms is recorded. That's still plenty to find patterns, because slow queries repeat. The same query shape running a thousand times an hour will show up in any reasonable sample.

Filtering

Recent versions let you replace the slowms threshold with an arbitrary filter over profile fields. This is the most precise way to capture only what you care about:

db.setProfilingLevel(1, {
  filter: {
    $or: [
      { millis: { $gte: 200 } },
      { docsExamined: { $gte: 10000 } },
      { planSummary: "COLLSCAN" },
    ],
  },
});

This records operations that are slow, or that examine lots of documents, or that do a collection scan, whether or not they're slow today. That last condition is especially useful: a collection scan on a small collection is fast now and becomes your next incident as the collection grows. When a filter is set, it replaces slowms and sampleRate for deciding what gets profiled.

To remove a filter, call db.setProfilingLevel(1, { filter: "unset" }).

Reading a Profile Document

Profiled operations land in the system.profile collection of the same database. Here's a typical entry, trimmed a little:

{
  op: "query",
  ns: "shop.orders",
  command: {
    find: "orders",
    filter: { status: "pending", region: "eu-west" },
    sort: { createdAt: 1 },
    limit: 100,
    $db: "shop"
  },
  keysExamined: 0,
  docsExamined: 1842210,
  hasSortStage: true,
  nreturned: 100,
  planSummary: "COLLSCAN",
  planCacheShapeHash: "5C1E2A7B",
  millis: 1270,
  numYield: 1439,
  locks: { /* ... */ },
  responseLength: 48211,
  ts: ISODate("2026-09-22T08:14:07.311Z"),
  client: "10.0.3.22",
  appName: "fulfillment-worker",
  user: "fulfillment@admin"
}

The fields that matter most:

  • op and ns: the operation type (query, update, remove, insert, command, getmore) and the namespace.
  • command: the full command, including the filter, sort, and projection. This is what you'll copy into explain().
  • keysExamined, docsExamined, nreturned: the same work-versus-result ratio that explain() reports. Here, 1.8 million documents examined to return 100.
  • planSummary: a one-line plan description like COLLSCAN or IXSCAN { status: 1, createdAt: 1 }. The fastest way to spot missing indexes.
  • hasSortStage: true means the results were sorted in memory, so the index didn't provide the order.
  • millis: total time in milliseconds.
  • planCacheShapeHash (called queryHash in older versions): identifies the query shape, so you can group many executions of the same query with different values.
  • appName, client, user: who sent it. This is why setting appName on every driver client pays off.

Queries That Find the Worst Offenders

system.profile is a normal (capped) collection, so you query it with the usual tools.

The most recent slow operations:

db.system.profile
  .find({}, { op: 1, ns: 1, millis: 1, planSummary: 1, appName: 1, ts: 1 })
  .sort({ ts: -1 })
  .limit(10);

Every collection scan recorded:

db.system.profile.find(
  { planSummary: "COLLSCAN" },
  { ns: 1, "command.filter": 1, millis: 1 },
);

Operations that examine far more than they return:

db.system.profile.find({
  nreturned: { $gt: 0 },
  $expr: { $gt: [{ $divide: ["$docsExamined", "$nreturned"] }, 1000] },
});

The most useful query of all groups executions by shape, so you see which query patterns cost the most in total, not just which single execution was slowest:

db.system.profile.aggregate([
  { $match: { op: { $in: ["query", "update", "remove", "command"] } } },
  {
    $group: {
      _id: { ns: "$ns", shape: "$planCacheShapeHash", plan: "$planSummary" },
      count: { $sum: 1 },
      totalMs: { $sum: "$millis" },
      avgMs: { $avg: "$millis" },
      maxMs: { $max: "$millis" },
      avgDocsExamined: { $avg: "$docsExamined" },
      sample: { $first: "$command" },
    },
  },
  { $sort: { totalMs: -1 } },
  { $limit: 10 },
]);
[
  {
    _id: { ns: "shop.orders", shape: "5C1E2A7B", plan: "COLLSCAN" },
    count: 412,
    totalMs: 498320,
    avgMs: 1209.5,
    maxMs: 2210,
    avgDocsExamined: 1839002,
    sample: {
      find: "orders",
      filter: { status: "pending", region: "eu-west" } /* ... */,
    },
  },
  {
    _id: { ns: "shop.products", shape: "A07F19C3", plan: "IXSCAN { tags: 1 }" },
    count: 3108,
    totalMs: 214400,
    avgMs: 69,
    maxMs: 180,
    avgDocsExamined: 5210,
    sample: {
      find: "products",
      filter: { tags: "sale", price: { $lte: 30 } } /* ... */,
    },
  },
];

This ranking is how you prioritize. The first shape is an obvious missing index. The second is subtler: each run is only 69 ms, but it runs thousands of times, and it examines 5,000 documents per call because the index covers tags but not price. A compound index would cut total load significantly. Sorting by totalMs surfaces both kinds of problem.

If you're profiling with slowms rather than a filter, remember you're only seeing operations above the threshold. A fast query that runs a million times an hour won't appear at all, even if it dominates your CPU. Temporarily lowering slowms or adding a docsExamined condition to a filter helps catch those.

Once you have a candidate, take the sample command and run it through explain("executionStats") to design the fix. Using explain() to Analyze and Debug Slow MongoDB Queries walks through that part.

Sizing system.profile

system.profile is a capped collection, 1 MB by default. On a busy database that's often only a few minutes of history, and old entries are overwritten as new ones arrive. To keep more, recreate it larger. You have to turn the profiler off first:

db.setProfilingLevel(0);
db.system.profile.drop();
db.createCollection("system.profile", { capped: true, size: 64 * 1024 * 1024 });
db.setProfilingLevel(1, { slowms: 50 });

64 MB is a reasonable size for an investigation. There's rarely a reason to go much larger; if you need long-term history, export the interesting entries or rely on a monitoring tool that keeps its own.

Making Settings Permanent

setProfilingLevel changes vanish on restart. For a baseline that always applies, use the operationProfiling section of the mongod config file:

# /etc/mongod.conf
operationProfiling:
  mode: slowOp # off | slowOp | all
  slowOpThresholdMs: 100
  slowOpSampleRate: 1.0

mode: slowOp is level 1 for every database on the instance. Many teams leave the mode off and keep only the threshold, which gives them slow query logging with no profiler writes, then turn profiling on for a specific database during an investigation.

The Slow Query Log

Even at profiling level 0, mongod writes every operation slower than slowms to its log as a structured JSON entry with the message "Slow query". It contains nearly the same information as a profile document, and it doesn't consume space in your database.

Because the log is JSON, jq makes it easy to rank slow queries straight from the file:

grep '"Slow query"' /var/log/mongodb/mongod.log \
  | jq -r '[.attr.ns, .attr.planSummary, .attr.durationMillis] | @tsv' \
  | sort -t$'\t' -k3 -nr \
  | head -20
shop.orders     COLLSCAN                        2210
shop.orders     COLLSCAN                        1984
shop.products   IXSCAN { tags: 1 }              180

The slow query log is often the better default for production. It's always on, costs almost nothing, and ships easily to whatever log platform you use (Loki, Elasticsearch, Datadog), where you can build dashboards and alerts on planSummary and durationMillis. The profiler is better for focused, interactive investigations where you want to query the data with MongoDB itself.

Replica Sets, Sharded Clusters, and Atlas

Replica sets. Profiling is configured per mongod, and system.profile isn't replicated. Enabling it on the primary tells you nothing about slow reads on secondaries. If your app reads from secondaries, enable profiling on those members too, connecting to each one directly.

Sharded clusters. On a mongos, setProfilingLevel only changes slowms and sampleRate for the router's own slow query log; there's no system.profile on mongos. To profile, enable it on each shard's mongod. The mongos log is still useful: it shows slow operations from the router's perspective, including time spent merging results from multiple shards.

Atlas. On Atlas, the Query Profiler in the cluster UI shows slow operations collected from the logs, with charts and per-shape grouping, and it doesn't require you to touch setProfilingLevel. The Performance Advisor analyzes slow queries and recommends indexes. On dedicated tiers, you can also run setProfilingLevel yourself when you want direct access to system.profile; free and Flex clusters have more restrictions on server-level settings, so rely on the Atlas UI there.

Keeping Overhead Under Control

Profiling isn't free. Every profiled operation causes an extra write to system.profile, and at level 2 that can double the write load on a busy database. Some guidelines:

  • Never leave level 2 on in production. Use it for minutes on a development or staging server, then turn it off.
  • Use level 1 with sampling or a filter for production investigations, and set a reminder to turn it off when you're done.
  • Prefer the slow query log for always-on monitoring. It's cheaper than profiler writes and survives restarts.
  • Watch for sensitive data. Profile documents and log entries include the full command, filter values and all. If your queries contain personal data, treat system.profile and your logs accordingly. MongoDB Enterprise and Atlas offer log redaction for this.
  • Don't set slowms to 0 on a busy server. Every operation hits the log, which can flood disk and make the log useless.

Common Pitfalls

Only looking at the slowest single query. The worst individual execution is often a one-off, like a report run once a day. The query shape with the highest total time is usually what's actually loading your server. Group by shape and sort by total.

Forgetting that slowms is instance-wide. Setting a low threshold in one database to investigate it also floods the log with slow operations from every other database on that server.

Profiling the primary when the problem is on a secondary. Profiling is per member. Check where the slow reads actually run.

Leaving the default 1 MB collection. You'll find that the entry you wanted was overwritten ten minutes ago. Size the collection before you start an investigation.

Missing appName. Without it, client gives you an IP address that could be any of dozens of pods. With it, you know which service to go fix.

Conclusion

The profiler and slow query log turn "the database feels slow" into a ranked list of query shapes with their cost, their plan, and the app that sent them. Start with the slow query log as your always-on baseline, turn on level 1 profiling with a filter when you need to dig in, group results by shape and total time, and hand the top offenders to explain() to design the fix.

Your next step: on a staging or production database, run db.setProfilingLevel(1, { filter: { $or: [{ millis: { $gte: 100 } }, { planSummary: "COLLSCAN" }] } }) for an hour of normal traffic, then run the grouping aggregation above. The top three rows are your performance roadmap for the week.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading