
MongoDB Time Series Collections: Storing IoT and Metrics Data Efficiently
Measurements arrive fast and never stop. A fleet of 5,000 sensors reporting every ten seconds produces over 43 million readings a day. Store each reading as its own document in a regular collection and you pay for it everywhere: every document carries its own field names, every reading adds an _id index entry, and a query for "average temperature per hour last week" has to touch millions of tiny documents.
The classic workaround was the bucket pattern, where you manually group readings into one document per sensor per hour. It works, but you have to write the bucketing logic, handle bucket overflows, and query the awkward nested shape. Time series collections do all of that for you. You insert one document per measurement, exactly as you'd like to, and MongoDB groups them into compressed, columnar buckets behind the scenes. Queries still see individual measurements.
This guide covers creating time series collections, choosing the timeField, metaField, and granularity, inserting and querying data, indexing, automatic expiry, updating and deleting measurements, migrating existing data, and the design mistakes that quietly ruin compression.
How Time Series Collections Work
When you create a time series collection, you tell MongoDB which field holds the timestamp (timeField) and, optionally, which field identifies the source of the data (metaField). MongoDB then stores measurements that share the same metaField value and fall into the same time window together in a single internal bucket document.
Inside a bucket, data is organized by column: all timestamps together, all temperature values together, and so on. Columns of similar values compress extremely well, and the metadata is stored once per bucket instead of once per reading. The result is typically a large reduction in storage and index size compared with one document per reading in a regular collection, along with much better cache efficiency.
You never interact with buckets directly. The collection you query is a view over an internal system.buckets.<name> collection, and MongoDB unpacks buckets as needed.
Creating a Time Series Collection
Time series collections must be created explicitly. You can't convert an existing collection in place.
db.createCollection("sensorReadings", {
timeseries: {
timeField: "ts",
metaField: "sensor",
granularity: "seconds",
},
expireAfterSeconds: 60 * 60 * 24 * 90, // keep 90 days
});
Each option shapes how data is bucketed:
timeField(required) is the field containing a BSON date for each measurement. Every document must have it, and it must be a real date, not a string or number.metaField(optional but strongly recommended) is the field that identifies the series, such as a device ID or a sub-document of tags. Measurements are only bucketed together if theirmetaFieldvalues are identical.granularitytells MongoDB how frequently data arrives per series, which determines bucket time spans.expireAfterSeconds(optional) deletes data older than the given age automatically.
Designing the metaField
The metaField is the most important design decision, because it controls which measurements share a bucket. It should contain values that identify the source and rarely or never change for that source:
{
ts: ISODate("2026-09-19T06:59:10Z"),
sensor: {
deviceId: "th-0042",
site: "warehouse-3",
type: "temp-humidity"
},
temperatureC: 21.4,
humidityPct: 48.2,
batteryV: 3.61
}
Here sensor is an object with identifying tags, and each measurement's readings sit at the top level. Every reading from th-0042 lands in the same series of buckets.
What you should not put in the metaField is anything that changes with every reading. If you put temperatureC inside sensor, every reading would have a unique metadata value, so every reading would get its own bucket, and you'd lose nearly all the benefit. The rule of thumb: metadata describes who sent the data, measurements describe what they sent.
A metaField with very high cardinality has a similar problem. If you have millions of distinct series that each produce only a few readings per bucket window, buckets stay small and compression suffers. Time series collections are at their best when each series produces a steady stream of data.
Choosing Granularity
Granularity should match how often each individual series reports, not the total ingestion rate across all devices:
| Granularity | Bucket span | Good for series reporting every... |
|---|---|---|
seconds | 1 hour | Few seconds to under a minute |
minutes | 24 hours | Minute or so |
hours | 30 days | Hour or more |
If your sensors report every 10 seconds, choose seconds. If a weather station reports every 15 minutes, minutes is closer. Picking a granularity that's too fine for your data creates many sparsely filled buckets; picking one that's too coarse can make queries over narrow time ranges unpack more data than needed.
In recent versions (MongoDB 6.3 and later), you can instead set custom bucketing parameters with bucketMaxSpanSeconds and bucketRoundingSeconds, which must be set to the same value:
db.createCollection("gridFrequency", {
timeseries: {
timeField: "ts",
metaField: "meter",
bucketMaxSpanSeconds: 300,
bucketRoundingSeconds: 300,
},
});
This is useful when none of the presets fit, for example when you always query in five-minute windows.
Inserting Measurements
You insert time series data like any other data, one document per measurement. Batch your inserts for throughput. Here's a Python ingester using PyMongo:
from datetime import datetime, timezone
import random
from pymongo import MongoClient
client = MongoClient("mongodb://localhost:27017")
readings = client["iot"]["sensorReadings"]
def read_sensors():
now = datetime.now(timezone.utc)
return [
{
"ts": now,
"sensor": {"deviceId": f"th-{i:04d}", "site": "warehouse-3", "type": "temp-humidity"},
"temperatureC": round(random.uniform(18, 26), 2),
"humidityPct": round(random.uniform(35, 60), 1),
}
for i in range(500)
]
batch = read_sensors()
result = readings.insert_many(batch, ordered=False)
print(f"inserted {len(result.inserted_ids)} readings")
ordered=False lets the server continue past individual failures and parallelize the batch, which is what you want for ingestion. Batches of a few hundred to a few thousand documents are a good starting point. The post on bulk write operations covers tuning batch sizes.
In Node.js, it's the same idea:
import { MongoClient } from "mongodb";
const client = new MongoClient(process.env.MONGODB_URI);
const readings = client.db("iot").collection("sensorReadings");
export async function ingest(messages) {
const docs = messages.map((m) => ({
ts: new Date(m.timestamp),
sensor: { deviceId: m.deviceId, site: m.site, type: m.type },
temperatureC: m.t,
humidityPct: m.h,
}));
await readings.insertMany(docs, { ordered: false });
}
Always convert timestamps to real Date objects. A string timestamp is rejected, since the timeField must be a BSON date.
Data doesn't need to arrive in perfect order. Late or out-of-order readings are fine, though heavily out-of-order data can lead to more, smaller buckets.
Querying Time Series Data
Queries look exactly like queries on a regular collection. You filter on the time field and metadata, and MongoDB uses bucket-level min and max values to skip buckets that can't match.
The latest reading from one device:
db.sensorReadings
.find({ "sensor.deviceId": "th-0042" })
.sort({ ts: -1 })
.limit(1);
Readings from the last hour for a site:
db.sensorReadings.find({
"sensor.site": "warehouse-3",
ts: { $gte: new Date(Date.now() - 60 * 60 * 1000) },
});
Downsampling With $dateTrunc
The most common analytics query rolls raw readings up into time buckets. $dateTrunc makes this clean:
db.sensorReadings.aggregate([
{
$match: {
"sensor.site": "warehouse-3",
ts: {
$gte: ISODate("2026-09-18T00:00:00Z"),
$lt: ISODate("2026-09-19T00:00:00Z"),
},
},
},
{
$group: {
_id: {
deviceId: "$sensor.deviceId",
hour: { $dateTrunc: { date: "$ts", unit: "hour" } },
},
avgTemp: { $avg: "$temperatureC" },
maxTemp: { $max: "$temperatureC" },
readings: { $sum: 1 },
},
},
{ $sort: { "_id.deviceId": 1, "_id.hour": 1 } },
]);
[
{
_id: { deviceId: "th-0042", hour: ISODate("2026-09-18T00:00:00Z") },
avgTemp: 21.7,
maxTemp: 22.9,
readings: 360,
},
{
_id: { deviceId: "th-0042", hour: ISODate("2026-09-18T01:00:00Z") },
avgTemp: 21.5,
maxTemp: 22.4,
readings: 360,
},
// ...
];
Change unit to "minute" with binSize: 15 for 15-minute buckets, or to "day" for daily summaries. $dateTrunc also accepts a timezone so daily buckets align with local midnight, which matters for dashboards. See handling dates and time zones for more on that.
Moving Averages and Gaps With Window Functions
$setWindowFields handles rolling calculations like moving averages, and $densify and $fill fill in missing intervals:
db.sensorReadings.aggregate([
{
$match: {
"sensor.deviceId": "th-0042",
ts: { $gte: ISODate("2026-09-18T00:00:00Z") },
},
},
{
$setWindowFields: {
partitionBy: "$sensor.deviceId",
sortBy: { ts: 1 },
output: {
tempMovingAvg: {
$avg: "$temperatureC",
window: { range: [-10, 0], unit: "minute" },
},
},
},
},
{ $project: { _id: 0, ts: 1, temperatureC: 1, tempMovingAvg: 1 } },
]);
The window uses a time range rather than a document count, so it stays correct even if some readings are missing. MongoDB 8.0 introduced block processing for many time series aggregations, which can make pipelines like these significantly faster without any changes to your queries.
Indexing Time Series Collections
For collections created in MongoDB 6.3 and later, MongoDB automatically creates a compound index on the metaField and timeField. That covers the most common query shape: "this series, this time range."
You can add secondary indexes for other access patterns, including on metadata subfields and measurement fields:
db.sensorReadings.createIndex({ "sensor.site": 1, ts: -1 });
db.sensorReadings.createIndex({ temperatureC: 1 });
The second index would help a query like "find any reading above 40 degrees" without scanning everything. Keep secondary indexes to what your queries actually need, since each one adds write cost. A 2dsphere index on a GeoJSON measurement field is supported too, which is handy for vehicle tracking.
Automatic Expiry
The expireAfterSeconds option on the collection deletes old data automatically. Unlike a TTL index on a regular collection, it removes whole buckets once every measurement in them is older than the threshold, which is far cheaper than deleting documents one at a time.
You can change it later with collMod:
db.runCommand({
collMod: "sensorReadings",
expireAfterSeconds: 60 * 60 * 24 * 30,
});
Because deletion happens per bucket, data can linger slightly past the threshold until the newest reading in its bucket is also old enough. If you need an exact cutoff in results, filter on ts in your queries. For regular collections, TTL indexes are the equivalent tool.
Updating and Deleting Measurements
Time series collections are optimized for append-only workloads. Early versions heavily restricted updates and deletes, allowing only operations that matched on the metaField. Recent versions have relaxed most of those restrictions, so you can delete by time range or update measurement fields, but these operations are still more expensive than on regular collections because MongoDB has to unpack, modify, and repack buckets.
Metadata updates are the efficient case, because they apply to whole buckets:
db.sensorReadings.updateMany(
{ "sensor.deviceId": "th-0042" },
{ $set: { "sensor.site": "warehouse-4" } },
);
Deleting a device's data is similarly efficient:
db.sensorReadings.deleteMany({ "sensor.deviceId": "th-0099" });
If your workload regularly corrects individual readings, keep that in mind. An occasional fix is fine; a pipeline that updates every reading after the fact is fighting the design.
Changing Settings After Creation
Some settings are fixed and some can change:
| Setting | Can change later? |
|---|---|
timeField | No |
metaField | No |
granularity | Only to a coarser value (seconds to minutes to hours) |
bucketMaxSpanSeconds / rounding | Can be increased |
expireAfterSeconds | Yes, with collMod |
Changing granularity uses collMod:
db.runCommand({
collMod: "sensorReadings",
timeseries: { granularity: "minutes" },
});
Because the timeField and metaField can't change, spend a few minutes getting the document shape right before you start ingesting. If you get it wrong, you'll need to create a new collection and copy the data.
Migrating Existing Data
If you already have readings in a regular collection, copy them into a new time series collection. In MongoDB 7.0.3 and later, $out can write directly into a time series collection:
db.legacyReadings.aggregate([
{
$project: {
_id: 0,
ts: { $toDate: "$timestamp" },
sensor: { deviceId: "$deviceId", site: "$site" },
temperatureC: "$temp",
humidityPct: "$humidity",
},
},
{
$out: {
db: "iot",
coll: "sensorReadingsTs",
timeseries: {
timeField: "ts",
metaField: "sensor",
granularity: "seconds",
},
},
},
]);
For very large migrations, sort the source by metadata and time and copy in batches, or use mongodump and mongorestore to move data first and reshape afterward. Sorted input produces fuller buckets and better compression. Compare the storage sizes afterward with db.sensorReadingsTs.stats(); the difference is often striking.
When Not to Use Time Series Collections
Time series collections aren't the right choice for everything with a timestamp:
- Documents that change frequently after insert, like orders or user profiles. These are entities, not measurements.
- Sparse, irregular events with unique metadata, where each series has only a handful of documents. Buckets won't fill and you won't see compression gains.
- Workloads that need features time series collections don't support, such as transactions writing into them, change streams on the collection, or unique indexes. Check the current restrictions for your version.
For event logs that are append-only and time-ordered, time series collections can still work well if the event source makes a sensible metaField. For a small rolling log that just needs a size cap, a capped collection may be simpler.
Common Pitfalls
Putting changing values in the metaField. Every unique metadata value starts a new bucket. Keep readings out of the metaField, and keep it to stable identifying tags.
Storing timestamps as strings. The timeField must be a BSON date. Convert at ingestion time.
Choosing granularity from total throughput. Granularity is about how often each series reports, not how many writes the cluster handles per second.
Treating it like a regular collection for updates. Frequent per-reading updates are expensive. Model corrections as new readings, or fix data before inserting.
Tiny single-document inserts. Sending one reading per request wastes round trips. Buffer and insert in batches with ordered: false.
Forgetting to filter on time. Queries without a time range have to consider every bucket. Almost every time series query should bound ts.
Conclusion
Time series collections let you store measurements the natural way, one document per reading, while MongoDB handles bucketing, columnar compression, and bulk expiry behind the scenes. The key decisions are made at creation time: a real date in the timeField, a stable identifying object in the metaField, and a granularity that matches how often each series reports. After that, queries are ordinary MongoDB queries, and aggregation tools like $dateTrunc and $setWindowFields handle downsampling and rolling statistics.
If you have a regular collection of readings today, create a time series collection next to it, copy a week of data over with $out, and compare stats() for both. The storage numbers will tell you whether it's worth migrating the rest.


