
MongoDB Atlas Monitoring and Alerts: What to Watch
One of the biggest reasons teams move to MongoDB Atlas is that they no longer have to build monitoring themselves. Metrics are collected automatically, charts are a click away, and alerts can reach Slack or PagerDuty in minutes. But "it's built in" isn't the same as "it's set up well." Plenty of Atlas projects run for years on default alerts, with nobody quite sure what the Metrics tab is telling them, until an incident forces a crash course.
The challenge isn't access to data. Atlas exposes more metrics than most people will ever look at. The challenge is knowing which few signals predict real problems, what healthy values look like, and how to configure alerts so they fire early for things that matter and stay quiet for things that don't.
This guide covers Atlas's monitoring tools, the metrics worth watching and why, how to configure alerts through the UI and the Atlas CLI, where to send notifications, and a practical baseline alert set you can adapt for your own clusters.
Atlas Monitoring Tools at a Glance
Atlas gives you several views, each suited to a different question:
| Tool | Best for |
|---|---|
| Metrics | Historical charts per node: CPU, disk, memory, operations, cache |
| Real-Time Performance Panel | Live view of operations, hottest collections, slowest queries |
| Query Profiler / Query Insights | Finding slow operations and their shapes over time |
| Performance Advisor | Index suggestions and index removal recommendations |
| Namespace Insights | Latency and operation counts per collection |
| Alerts | Notifications when metrics or events cross thresholds |
Availability varies by tier. The free M0 cluster and Flex clusters include a limited set of metrics and alerts. Dedicated clusters (M10 and up) get the full toolset, including the Real-Time Performance Panel, the Performance Advisor, and finer-grained metric resolution. If you're running production workloads, you're almost certainly on a dedicated tier, and everything below applies.
Reading the Metrics Tab
Open your cluster, choose Metrics, and you'll see charts for each node in the replica set. Two habits make this view much more useful.
First, select the primary and a secondary separately rather than looking at an aggregate. Write-heavy problems show up on the primary. Replication and read-preference problems show up on secondaries.
Second, zoom out before you zoom in. A CPU chart at 70% means nothing on its own. Compare it to the same hour last week. Atlas lets you change the granularity and time range, and a 30-day view is often where slow-burning trends (disk growth, rising connections, a gradually filling cache) become obvious.
The Metrics Worth Watching
Normalized CPU
Atlas shows both raw and normalized CPU. Normalized values are scaled to the number of cores, so 100% means the node is fully busy regardless of instance size. Use normalized CPU for alerts; raw values change meaning every time you scale.
Also watch CPU steal on the burstable, lower dedicated tiers. Some small instance types use CPU credits. When credits run out, the provider throttles the instance, and steal time climbs. Sustained steal means you've outgrown the tier.
Healthy: bursts to 70% or 80% are fine. Sustained above 80% means queries will queue at peak and you should find the expensive operations (usually via the Profiler or Performance Advisor) or scale up.
Query Targeting
Query targeting is the ratio of documents (or index keys) scanned to documents returned. A ratio of 1 means every document examined was returned: perfectly targeted. A ratio of 1,000 means the server examined a thousand documents for every one it sent back.
This is one of the most valuable metrics in Atlas because it catches missing or poor indexes before they cause a CPU crisis. Atlas includes a default alert for high query targeting. When it fires, go to the Performance Advisor or Query Insights to find the offending query shape, then check it with explain():
db.orders
.find({ status: "pending", region: "eu-west" })
.sort({ createdAt: -1 })
.explain("executionStats").executionStats;
{
nReturned: 42,
totalKeysExamined: 0,
totalDocsExamined: 318440,
executionTimeMillis: 612
}
That's a collection scan with a targeting ratio in the thousands. An index on { status: 1, region: 1, createdAt: -1 } would bring it close to 1.
Disk IOPS, Latency, and Space
Disk metrics tell you whether storage is keeping up:
- Disk IOPS compared to the provisioned IOPS for your cluster. Hitting the ceiling means operations wait on disk.
- Disk latency (read and write). Sustained increases usually mean the working set no longer fits in memory, or IOPS are exhausted.
- Disk space used (percent). Atlas can auto-scale storage if you enable it, but alerts still matter: auto-scaling has limits, and rapid unexpected growth is itself a signal.
Cache Usage and Memory
Under the hardware metrics, look at WiredTiger cache activity: cache used, dirty bytes, and bytes read into cache. A steady, high rate of bytes read into cache means the working set is spilling out of memory, which is the classic precursor to rising disk latency and slower queries. The mechanics are explained in MongoDB Performance Tuning: Memory, WiredTiger Cache, and Disk I/O.
Connections
Every tier has a maximum number of connections. Atlas charts current connections and offers an alert on connections as a percentage of the limit. A sudden jump usually follows a deploy that creates a new client per request or runs many serverless function instances, each with its own pool.
Replication Lag and Oplog Window
Replication lag is how far behind the primary a secondary is, in seconds. A few seconds during write bursts is normal. Minutes of lag mean secondaries can't keep up, which puts majority writes at risk of slowing down and makes secondary reads stale.
The replication oplog window is how many hours of operations the oplog holds. If a secondary's lag ever exceeds the window, it falls off the oplog and needs a full initial sync. Watch for the window shrinking during heavy write periods, such as nightly batch jobs or backfills.
Operation Execution Time and Opcounters
Opcounters (operations per second by type) are your baseline for load. Operation execution time shows average latency for reads, writes, and commands. Read the two together: latency up with flat opcounters points to a regression (a new query, a dropped index, data growth). Latency up with opcounters doubling points to capacity.
Tickets and Queues
When the storage engine is saturated, operations queue for execution tickets. Atlas charts tickets available and queues. Queues above zero for more than a moment indicate the node is overloaded, usually as a result of one of the problems above.
Metrics You Can Mostly Ignore
Not every chart deserves an alert. Page faults in isolation, raw (non-normalized) CPU, network bytes without context, and individual opcounter types rarely need a pager. They're useful when investigating, not as triggers. Alerting on them mostly teaches people to ignore alerts.
Configuring Alerts
Alerts in Atlas are configured at the project level, under Project > Alerts > Alert Settings. Each alert has a condition (a metric threshold or an event), a scope (which clusters or hosts), and one or more notification targets.
New projects come with a set of default alerts, including host down, replica set without a primary, high query targeting, disk space, and connections. Review them rather than accepting them as-is. The defaults are a floor, not a strategy.
Condition Types
Atlas alerts fall into two broad categories:
- Metric thresholds, for example "Normalized System CPU is above 80% for 10 minutes."
- Events, for example "Replica set elected a new primary," "Backup snapshot failed," "User joined the project," or "Cluster auto-scaling triggered."
Events are underrated. An unexpected primary election, a failed backup, or a change to network access can matter as much as any metric.
Creating Alerts with the Atlas CLI
Clicking through the UI works for one project. For many projects, or to keep alert configuration in version control, use the Atlas CLI. Log in once:
atlas auth login
List existing alert configurations:
atlas alerts settings list --projectId 64f1c0a2e4b0a1b2c3d4e5f6 --output json
Create a metric-threshold alert. The exact flag names can vary between CLI versions, so check atlas alerts settings create --help on your installation:
atlas alerts settings create \
--projectId 64f1c0a2e4b0a1b2c3d4e5f6 \
--event OUTSIDE_METRIC_THRESHOLD \
--metricName NORMALIZED_SYSTEM_CPU_USER \
--metricOperator GREATER_THAN \
--metricThreshold 80 \
--metricUnits RAW \
--metricMode AVERAGE \
--notificationType EMAIL \
--notificationEmailAddress oncall@example.com \
--notificationIntervalMin 60 \
--enabled
For anything more complex than a single notification target, recent versions of the CLI can also create alert configurations from a JSON file (via a --file flag) matching the Atlas Administration API schema:
{
"eventTypeName": "OUTSIDE_METRIC_THRESHOLD",
"enabled": true,
"metricThreshold": {
"metricName": "OPLOG_MASTER_TIME",
"operator": "LESS_THAN",
"threshold": 24,
"units": "HOURS",
"mode": "AVERAGE"
},
"notifications": [
{
"typeName": "GROUP",
"roles": ["GROUP_OWNER"],
"intervalMin": 60,
"delayMin": 0,
"emailEnabled": true
}
]
}
Keeping these files in a repository means every project gets the same alert baseline, and changes are reviewed like code. If you manage Atlas with Terraform, the MongoDB Atlas provider exposes alert configurations as resources, which works just as well.
Where to Send Alerts
Atlas integrates with the tools most teams already use: email, SMS, Slack, PagerDuty, Opsgenie, Microsoft Teams, Datadog, and generic webhooks, among others. Configure integrations under Project > Integrations, then choose them as notification targets on individual alerts.
A reasonable routing pattern:
- Critical (host down, no primary, disk nearly full, backup failed): PagerDuty or Opsgenie, so someone is paged.
- Warning (CPU sustained high, replication lag, connections above 80%, query targeting): a team Slack or Teams channel, reviewed during working hours.
- Informational (elections, auto-scaling events, configuration changes): a low-traffic channel or email digest for audit purposes.
Use the delay and interval settings. A delay of a few minutes filters out blips (like a brief CPU spike during a snapshot), and a reasonable re-notification interval prevents the same alert from flooding a channel.
Webhooks are handy when you want alerts in a system Atlas doesn't integrate with directly. Atlas sends a JSON payload to your endpoint; secure it with a secret and verify the request signature on your side.
A Practical Baseline Alert Set
Here's a starting point for a production dedicated cluster. Tune thresholds to your workload after a few weeks of observing normal behavior.
| Alert | Threshold | Severity |
|---|---|---|
| Host down | Any | Critical |
| Replica set has no primary | Any | Critical |
| Disk space used | Above 85% | Critical |
| Backup snapshot failed or delayed | Any | Critical |
| Normalized system CPU | Above 80% for 15 min | Warning |
| Replication lag | Above 60 s for 5 min | Warning |
| Replication oplog window | Below 24 hours | Warning |
| Connections (% of limit) | Above 80% | Warning |
| Query targeting (scanned / returned) | Above 1000 | Warning |
| Disk read or write latency | Above your baseline for 10 min | Warning |
| Primary election | Any | Info |
| Cluster auto-scaling event | Any | Info |
Two principles behind this list: every critical alert should require action from a human right now, and every warning should be something you'd want fixed within a day. If an alert doesn't meet either bar, lower its severity or delete it.
Using Monitoring During an Incident
When an alert fires, a consistent routine saves time:
- Check scope. Is it one node or all of them? The primary or a secondary? One cluster or the whole project?
- Look at the Real-Time Performance Panel. It shows current operations and the hottest collections, which often points directly at the culprit.
- Correlate with deploys. Most regressions follow a change. Compare the metric timeline to your deploy history.
- Find the query. Use Query Insights or the Profiler to find slow operations that started around the same time, then check the Performance Advisor for index suggestions.
- Decide: fix or scale. A bad query needs an index or a code fix. Genuine growth needs a larger tier or sharding. Atlas makes scaling up easy, but scaling to hide a missing index just makes the bill bigger.
Common Mistakes
Relying only on default alerts. The defaults don't cover everything that matters for your workload, like oplog window or backup failures in some projects. Review them explicitly.
Alerting on everything. A channel that fires dozens of alerts a day gets muted. Keep critical alerts rare and meaningful.
Using raw CPU for thresholds. Raw CPU changes meaning when you change instance size. Use normalized CPU.
Ignoring events. Unexpected elections, failed snapshots, and network access changes are often more important than metric blips.
Never revisiting thresholds. A threshold that made sense at launch may be far too loose or too tight a year later. Review alert history quarterly and adjust.
Conclusion
Atlas does the hard work of collecting metrics, but you still have to decide what matters. Focus on a small set of signals: normalized CPU, query targeting, disk IOPS and latency, cache activity, connections, replication lag, and oplog window. Configure alerts at the project level with clear severities, route them to the right places, and keep the configuration in code with the Atlas CLI or Terraform so every project gets the same baseline.
Your next step: open Alert Settings in your production project today, compare it to the baseline table above, and add the two alerts most teams are missing: oplog window below 24 hours and failed backup snapshots.


