
Monitoring MongoDB with Prometheus and Grafana
Most MongoDB incidents don't arrive out of nowhere. Replication lag creeps up for an hour before a secondary falls off the oplog. Connections climb steadily after a deploy until the server refuses new ones. The WiredTiger cache fills a little more every week until one Monday morning, latency doubles. The warning signs were there. Nobody was watching them.
If you run MongoDB yourself, on VMs, bare metal, or Kubernetes, Prometheus and Grafana are the most common open-source way to watch it. Prometheus scrapes and stores time-series metrics, Grafana turns them into dashboards, and Alertmanager routes alerts to Slack, PagerDuty, or email. MongoDB doesn't speak Prometheus natively, so an exporter sits in between and translates server statistics into metrics.
This guide covers setting up the exporter with the right permissions, running the full stack with Docker Compose, the metrics worth graphing, useful PromQL queries, and alert rules that catch real problems without paging you for noise.
How the Pieces Fit Together
The flow is straightforward:
- MongoDB exposes its internal statistics through commands like
serverStatus,replSetGetStatus, anddbStats. - The exporter connects to MongoDB as a regular client, runs those commands on each scrape, and exposes the results as Prometheus metrics over HTTP (by default on port 9216).
- Prometheus scrapes the exporter's
/metricsendpoint on an interval and stores the time series. - Grafana queries Prometheus to draw dashboards.
- Alertmanager receives alerts fired by Prometheus rules and routes them to people.
The de facto standard exporter is the open-source Percona MongoDB Exporter (percona/mongodb_exporter). It's actively maintained, supports replica sets and sharded clusters, and powers Percona's own monitoring product.
Run one exporter per mongod (and per mongos if you want router metrics). A single exporter pointed at a replica set connection string only sees the member it's connected to, which hides exactly the per-node problems you want to catch.
Step 1: Create a Monitoring User
The exporter needs read access to server statistics, not to your data. The built-in clusterMonitor role covers serverStatus, replica set status, and similar commands. Add read on the local database so the exporter can inspect the oplog:
db.getSiblingDB("admin").createUser({
user: "exporter",
pwd: passwordPrompt(),
roles: [
{ role: "clusterMonitor", db: "admin" },
{ role: "read", db: "local" },
],
});
If you want per-collection or per-index metrics (which the exporter can collect with extra flags), the user also needs read access to those databases. Start without them. Collection-level metrics multiply the number of time series quickly, and on a database with thousands of collections they can overwhelm Prometheus.
Step 2: Run the Stack with Docker Compose
Here's a minimal, working stack for a local or staging environment. It runs a single-node MongoDB replica set, the exporter, Prometheus, and Grafana.
# docker-compose.yml
services:
mongo:
image: mongo:8.0
command: ["--replSet", "rs0", "--bind_ip_all"]
ports:
- "27017:27017"
volumes:
- mongo-data:/data/db
mongodb-exporter:
image: percona/mongodb_exporter:0.43 # pin to the current release
command:
- "--mongodb.uri=mongodb://exporter:${EXPORTER_PASSWORD}@mongo:27017/admin?directConnection=true"
- "--collect-all"
- "--compatible-mode"
ports:
- "9216:9216"
depends_on:
- mongo
prometheus:
image: prom/prometheus:latest
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
- ./alerts.yml:/etc/prometheus/alerts.yml:ro
- prom-data:/prometheus
ports:
- "9090:9090"
grafana:
image: grafana/grafana:latest
volumes:
- ./grafana/provisioning:/etc/grafana/provisioning:ro
- grafana-data:/var/lib/grafana
ports:
- "3000:3000"
volumes:
mongo-data:
prom-data:
grafana-data:
A few notes on the exporter flags:
--mongodb.uriusesdirectConnection=trueso the exporter monitors exactly this node instead of following the replica set to the primary.--collect-allenables the optional collectors (database stats, collection stats, index stats, top metrics, and so on). On large deployments, enable only the collectors you need rather than all of them.--compatible-modeadditionally exposes metrics using the naming scheme of the older exporter generation, which many community dashboards still expect.
For production, pin every image tag (including Prometheus and Grafana) instead of using latest, and pass the password through a secret rather than an environment variable when your platform supports it. If you're new to running MongoDB in containers, Running MongoDB in Docker and Docker Compose covers initializing the replica set and creating users.
Prometheus Configuration
# prometheus.yml
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- /etc/prometheus/alerts.yml
scrape_configs:
- job_name: mongodb
static_configs:
- targets: ["mongodb-exporter:9216"]
labels:
cluster: staging-rs0
node: mongo-1
In production you'd list one exporter per node, or use service discovery (Kubernetes, Consul, EC2) to find them. Adding cluster and node labels up front makes every query and alert easier to read later.
Grafana Data Source Provisioning
Provisioning the data source means Grafana is ready the moment it starts:
# grafana/provisioning/datasources/prometheus.yml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
Start everything and confirm the exporter can reach MongoDB:
docker compose up -d
curl -s localhost:9216/metrics | grep -E "^mongodb_up"
mongodb_up 1
mongodb_up 1 means the exporter connected and ran its commands. A value of 0 almost always means a wrong URI, bad credentials, or a missing role.
Step 3: Explore the Metrics
The exporter exposes hundreds of metrics. Before building dashboards, look at what's actually there, because metric names depend on the exporter version and whether compatible mode is on:
curl -s localhost:9216/metrics | grep -v "^#" | cut -d"{" -f1 | cut -d" " -f1 | sort -u | head -40
In the exporter's newer naming scheme, most metrics mirror the structure of serverStatus. For example, serverStatus().opcounters becomes mongodb_ss_opcounters, and serverStatus().wiredTiger.cache fields become mongodb_ss_wt_cache_* metrics. Replica set status appears under mongodb_rs_*. Once you know that mapping, finding a metric for any serverStatus field is a matter of searching.
The PromQL examples below use the newer names. If your output differs, adjust the names to match what /metrics shows.
The Metrics That Matter
You don't need a hundred panels. A focused dashboard answers five questions: is it up, is it busy, is it healthy, is it keeping up, and is it running out of anything.
Availability
mongodb_up
The simplest and most important. Alert if it's 0, or if the target disappears from Prometheus entirely (up{job="mongodb"} == 0).
Operation Throughput
sum by (node, legacy_op_type) (rate(mongodb_ss_opcounters[5m]))
This gives inserts, queries, updates, deletes, getmores, and commands per second. It's your baseline for "normal." Most other anomalies make more sense when viewed next to throughput: latency rising with flat throughput is a very different problem from latency rising with doubled throughput.
Operation Latency
serverStatus().opLatencies tracks cumulative latency and operation counts for reads, writes, and commands. Dividing the rate of one by the rate of the other gives average latency:
rate(mongodb_ss_opLatencies_latency{op_type="reads"}[5m])
/
rate(mongodb_ss_opLatencies_ops{op_type="reads"}[5m])
The result is in microseconds. Averages hide tail latency, so treat this as a trend indicator. For percentile latency, measure in your application (driver command monitoring or APM), where you can use histograms.
Connections
mongodb_ss_connections{conn_type="current"}
Graph current connections alongside available. A steady climb after a deploy usually means an application isn't reusing its client or has an oversized pool. The fix lives in the app, as covered in MongoDB Connection Pooling: Avoiding Too Many Connections.
Replication Health
For replica sets, you want member states and lag. Member state is exposed per member (1 is PRIMARY, 2 is SECONDARY). Replication lag can be computed from each member's optime relative to the primary. Depending on the exporter version, lag may be available directly as a metric or you may compute it from optime timestamps. Check /metrics for mongodb_rs_members_* series on your version.
Also watch the oplog window: how many hours of operations the oplog currently holds. If a secondary falls further behind than the window, it can't catch up and needs a full resync. A shrinking window during heavy write periods is a signal to grow the oplog.
WiredTiger Cache
mongodb_ss_wt_cache_bytes_currently_in_the_cache
/ mongodb_ss_wt_cache_maximum_bytes_configured
mongodb_ss_wt_cache_tracked_dirty_bytes_in_the_cache
/ mongodb_ss_wt_cache_maximum_bytes_configured
Cache fill around 80% is normal. Sustained values near 95%, or dirty data near 20%, mean application threads are being drafted into eviction and latency will suffer. Pair these with the rate of pages read into cache, which shows how often the working set misses memory.
Queues and Tickets
When operations queue waiting for the storage engine, users feel it immediately. Graph mongodb_ss_globalLock_currentQueue (readers and writers) and, on newer servers, the execution ticket metrics. Any sustained queue is worth investigating.
Host Metrics
The exporter only knows what MongoDB reports. Run node_exporter on each database host too, and graph CPU, memory, disk latency, disk throughput, and free disk space. Disk space in particular deserves a dedicated alert, because a full disk takes a mongod down hard.
Building Dashboards
You have two good options:
- Import a community dashboard. Percona publishes MongoDB dashboards (originally built for Percona Monitoring and Management) that work with this exporter, and Grafana's dashboard library has several more. Import one to get started quickly, then prune what you don't use.
- Build a focused one yourself. Six to ten panels covering the metrics above, templated by
clusterandnodevariables, is often more useful during an incident than a wall of 60 graphs.
Whichever you choose, put the panels in the order you'd debug: availability and replication state at the top, then throughput and latency, then connections, cache, queues, and finally host resources. During an incident you'll scan top to bottom.
Export finished dashboards as JSON and commit them next to your provisioning files so they're versioned and reproducible:
# grafana/provisioning/dashboards/mongodb.yml
apiVersion: 1
providers:
- name: mongodb
folder: Databases
type: file
options:
path: /etc/grafana/provisioning/dashboards/json
Writing Alert Rules
Good alerts are specific, actionable, and have a for duration so they don't fire on a single bad scrape. Here's a starting set:
# alerts.yml
groups:
- name: mongodb
rules:
- alert: MongoDBDown
expr: mongodb_up == 0 or up{job="mongodb"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "MongoDB {{ $labels.node }} is unreachable"
- alert: MongoDBConnectionsHigh
expr: |
mongodb_ss_connections{conn_type="current"}
/ (mongodb_ss_connections{conn_type="current"}
+ mongodb_ss_connections{conn_type="available"}) > 0.8
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.node }} is using over 80% of available connections"
- alert: MongoDBCacheDirtyHigh
expr: |
mongodb_ss_wt_cache_tracked_dirty_bytes_in_the_cache
/ mongodb_ss_wt_cache_maximum_bytes_configured > 0.15
for: 10m
labels:
severity: warning
annotations:
summary: "WiredTiger dirty cache above 15% on {{ $labels.node }}"
- alert: MongoDBQueuedOperations
expr: sum by (node) (mongodb_ss_globalLock_currentQueue) > 20
for: 5m
labels:
severity: warning
annotations:
summary: "Operations are queuing on {{ $labels.node }}"
- alert: HostDiskSpaceLow
expr: |
node_filesystem_avail_bytes{mountpoint="/var/lib/mongo"}
/ node_filesystem_size_bytes{mountpoint="/var/lib/mongo"} < 0.15
for: 15m
labels:
severity: critical
annotations:
summary: "Less than 15% disk space left for MongoDB data"
Add a replication lag alert using whichever lag metric your exporter version provides, with a threshold that reflects your tolerance (for example, more than 30 seconds for 5 minutes).
Validate rules before reloading Prometheus:
docker compose exec prometheus promtool check rules /etc/prometheus/alerts.yml
Checking /etc/prometheus/alerts.yml
SUCCESS: 5 rules found
Common Pitfalls
Pointing one exporter at the whole replica set. The exporter follows the connection string to one member. Run one per node with directConnection=true so each member's health is visible.
Collecting everything on large deployments. Collection and index collectors create series per namespace. With thousands of collections, that's hundreds of thousands of series. Enable those collectors selectively or with a filtered list of namespaces.
Alerting on averages alone. Average latency looks fine while the p99 is on fire. Use server-side averages for trends and application-side histograms for latency SLOs.
No for durations. Without them, a single slow scrape or brief spike pages someone at 3 a.m. Most database alerts should need a condition to hold for at least a few minutes.
Forgetting host metrics. Many MongoDB problems are really disk or memory problems. Without node_exporter, you'll see the symptom and miss the cause.
Unpinned image versions. A new exporter release can rename metrics and silently break dashboards. Pin versions and upgrade intentionally.
What About MongoDB Atlas?
If you use Atlas, you already get built-in metrics, alerts, and the Performance Advisor, so you may not need any of this. Atlas also offers a Prometheus integration on dedicated clusters that lets your own Prometheus scrape Atlas metrics through service discovery, which is useful when you want database metrics in the same Grafana as the rest of your stack. MongoDB Atlas Monitoring and Alerts: What to Watch covers the native tooling.
Conclusion
Prometheus and Grafana give you durable, queryable visibility into self-managed MongoDB. A least-privilege monitoring user, one Percona exporter per node, a Prometheus scrape job with sensible labels, and a focused Grafana dashboard cover the essentials. Add alert rules for availability, connections, cache pressure, queues, replication, and disk space, each with a for duration, and you'll hear about problems while they're still small.
Your next step: spin up the Compose stack from this guide against a staging replica set, run the curl ... | grep command to list the metric names your exporter version actually produces, and build your first three panels from them: throughput, connections, and cache fill.


