Type something to search...
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's orders, today's sessions, the last hour of sensor readings. Everything older is rarely touched but still sits on the same high-performance storage, inflating your disk size, your backup size, and eventually your cluster tier. You end up paying premium prices to store data that gets queried once a quarter.

You could delete old data with a TTL index, but compliance, analytics, or customer support often need it. You could write an export job to S3, but then you have a pipeline to maintain and a second place to query. Atlas Online Archive is the managed middle ground. It automatically moves documents that match a rule from your cluster to MongoDB-managed cloud object storage, and it keeps them queryable through a single connection string alongside your live data.

This guide covers how Online Archive works, how to choose archiving rules and partition fields, how to query archived data, what it does to your costs, and the limitations you need to design around.

How Online Archive Works

You configure an archiving rule on a collection. The rule says which documents are eligible for archiving, usually "documents whose date field is older than N days." Atlas then runs a background job that:

  1. Finds documents matching the rule, using an index you provide.
  2. Copies them to cloud object storage managed by MongoDB, organized by the partition fields you chose.
  3. Deletes them from the cluster once the copy is confirmed.

Archived documents are stored in a compressed, columnar-friendly format and are exposed through Data Federation. Atlas gives you connection strings that query the archive alone, or the cluster and archive together as one logical collection.

A few properties are important to understand up front:

  • Online Archive is available on dedicated clusters (M10 and above), not on M0 or Flex.
  • Archived data is read-only. You can't update or delete individual archived documents through normal queries.
  • Archiving runs periodically in the background, so documents aren't moved the exact second they become eligible.
  • Your application's normal connection string only sees the cluster. To include archived data, you use the federated connection string.

Choosing What to Archive

Online Archive works best for data that's naturally append-mostly and time-ordered:

Good candidatesPoor candidates
Event logs and audit trailsUser profiles and accounts
Orders and invoices older than a yearDocuments that are updated long after creation
IoT readings and metricsSmall reference collections
Chat and notification historyData that must be queried with low latency forever

The key question is: once a document is old, is it effectively immutable? If old orders might still be refunded or edited, set the archive age past the point where that can happen, or use a custom rule based on status.

Also check whether the data should be kept at all. If nobody needs sensor readings older than 90 days, a TTL index is simpler and cheaper than archiving.

Date-Based Archiving Rules

The most common rule archives documents based on a date field. For an orders collection with a createdAt field, you might archive everything older than 180 days.

First, make sure the archiving job can find eligible documents efficiently. Create an index on the date field:

db.orders.createIndex({ createdAt: 1 });

Then, in the Atlas UI, open your cluster, go to the Online Archive tab, and click Configure Online Archive. Choose:

  • Namespace: shop.orders
  • Archiving rule: date match, field createdAt, archive after 180 days
  • Date format: ISODate (you can also use epoch seconds, milliseconds, or nanoseconds if your field stores a number)
  • Partition fields: covered in the next section

The date field must be a real date (or a numeric epoch value with the matching format). If your dates are stored as strings, archiving can't evaluate them, which is one more reason to store dates correctly.

Custom Criteria Rules

Sometimes age alone isn't the right rule. Maybe only closed support tickets should be archived, or only orders in a final state. A custom criteria rule lets you specify a query filter:

{
  "status": { "$in": ["delivered", "refunded", "cancelled"] },
  "updatedAt": { "$lt": { "$date": "2026-03-01T00:00:00Z" } }
}

Custom rules are powerful but easy to get wrong. A static date in a custom query doesn't move forward over time, so you'll have to update it periodically, and the archiving job needs an index that supports the filter:

db.orders.createIndex({ status: 1, updatedAt: 1 });

For most time-series-like data, the date-based rule is simpler and keeps working without maintenance.

Choosing Partition Fields

Archived data is organized in object storage by partition fields, and they have a huge effect on how fast (and how cheap) queries against the archive are. When you query the archive with a filter on a partition field, Atlas can skip every partition that doesn't match.

For date-based rules, the date field is always used as a partition. You can add a small number of additional partition fields (check the current limit in the docs; it's intentionally small), in order of how often your queries filter on them.

Think about how you'll actually query archived data:

  • Support looks up a customer's old orders: partition by customerId.
  • Finance runs monthly reports by region: partition by region.
  • Analytics filters by device type: partition by deviceType.

Order matters. Put the field you filter on most frequently first. A good partition field has moderate cardinality: region (a handful of values) or customerId (many values) both work, while a field that's unique per document or always the same value provides little benefit.

Partition fields can't be changed after the archive is created. If you pick the wrong ones, you'll need to create a new archive, so spend five minutes listing your real archive queries before you click save.

Configuring an Archive Through the API

For repeatable environments, configure archives with the Atlas Admin API or the Atlas CLI instead of the UI. The request body for a date-based archive looks roughly like this (check the Admin API reference for the exact current schema):

{
  "dbName": "shop",
  "collName": "orders",
  "criteria": {
    "type": "DATE",
    "dateField": "createdAt",
    "dateFormat": "ISODATE",
    "expireAfterDays": 180
  },
  "partitionFields": [
    { "fieldName": "createdAt", "order": 0 },
    { "fieldName": "customerId", "order": 1 },
    { "fieldName": "region", "order": 2 }
  ],
  "dataExpirationRule": {
    "expireAfterDays": 2555
  }
}

Note the naming: in the criteria, expireAfterDays means "archive documents older than this many days." The separate dataExpirationRule controls when archived data is permanently deleted from the archive, here roughly seven years. That's how you implement a full retention policy: hot in the cluster for six months, cold in the archive until seven years, then gone.

Keeping this configuration in source control, next to your index definitions, makes it much easier to review and reproduce across staging and production.

Querying Archived Data

After you create an archive, the cluster's Connect dialog offers additional connection strings:

  • Cluster only: your normal connection string. Archived documents are invisible.
  • Archive only: queries just the archived data.
  • Cluster and archive: a federated endpoint that returns results from both.

Connect with the combined string when you need the full history:

mongosh "mongodb://atlas-online-archive-abc123-xyz.a.query.mongodb.net/shop" \
  --tls --authenticationDatabase admin --username support-reader

Then query as normal:

db.orders
  .find(
    { customerId: ObjectId("66a1f0c2e4b0a1b2c3d4e5f6") },
    { _id: 1, createdAt: 1, total: 1, status: 1 },
  )
  .sort({ createdAt: -1 });

The results include both recent orders from the cluster and older ones from the archive, and because customerId is a partition field, the archive side only reads that customer's partitions.

In application code, keep two clients: one for normal operations and one for the rare historical queries.

import { MongoClient } from "mongodb";

const live = new MongoClient(process.env.MONGODB_URI);
const history = new MongoClient(process.env.MONGODB_ARCHIVE_URI);

export async function getOrderHistory(customerId) {
  return history
    .db("shop")
    .collection("orders")
    .find({ customerId })
    .sort({ createdAt: -1 })
    .limit(100)
    .toArray();
}

export async function getRecentOrders(customerId) {
  return live
    .db("shop")
    .collection("orders")
    .find({ customerId })
    .sort({ createdAt: -1 })
    .limit(20)
    .toArray();
}

This makes the performance characteristics explicit. Recent-order lookups hit indexes on the cluster. Full-history lookups go through federation and are expected to be slower.

What It Does to Your Bill

Online Archive changes where you pay, and for the right data it can reduce total cost significantly. The exact rates vary by cloud and region, so check the current Atlas pricing page, but the moving parts are:

What gets cheaper:

  • Cluster storage. Archived data leaves your cluster's disks, which are the most expensive storage you have.
  • Cluster tier. A smaller working set and smaller disks can let you drop a tier, which is often the biggest saving.
  • Backups. Snapshots only include data on the cluster, so backup storage shrinks with it.
  • Operational headroom. Index builds, initial syncs, and resyncs finish faster on a smaller dataset.

What you now pay for:

  • Archive storage, billed per GB per month at object-storage-like rates, far below cluster disk.
  • Data processed by queries against the archive, similar to Data Federation.
  • The archiving process itself, which reads from your cluster.

The savings are largest when the archive holds most of your data and gets queried rarely. If you'll run heavy analytics over archived data every hour, the query charges can add up, and a dedicated analytics pipeline might suit you better.

A useful sanity check before enabling an archive:

const cutoff = new Date(Date.now() - 180 * 24 * 60 * 60 * 1000);

const total = await db.collection("orders").estimatedDocumentCount();
const eligible = await db
  .collection("orders")
  .countDocuments({ createdAt: { $lt: cutoff } });

console.log(`${((eligible / total) * 100).toFixed(1)}% eligible to archive`);
78.4% eligible to archive

If most of the collection is eligible and those documents are rarely read, Online Archive is likely worth it.

Operational Considerations

Watch the first run. When you first enable an archive on a large collection, there may be a big backlog of eligible documents. Archiving that backlog adds read and delete load to the cluster. Enable it during a quieter period and monitor the cluster.

Disk space isn't reclaimed instantly. Deleting documents frees space inside WiredTiger's data files for reuse, but the files may not shrink immediately. Your storage metrics will reflect the change over time, and new writes reuse the freed space.

Pause when needed. You can pause and resume an archive. Pause it before bulk maintenance or migrations that might interact badly with archiving deletes.

Plan for the archive in restores. Restoring a cluster snapshot doesn't restore the archive, which is stored separately. Think through how archived data fits into your disaster recovery plan.

Sharded clusters. Online Archive supports sharded collections. Choose partition fields with your query patterns in mind, and note that the shard key and the archive partitions are separate concepts.

Common Pitfalls

Archiving data that still changes. Once a document is archived, you can't update it with normal operations. Set the archive age well past the point where edits, refunds, or late-arriving updates stop.

Missing index on the rule's fields. Without an index on the date field (or custom criteria fields), the archiving job has to scan the collection, putting unnecessary load on the cluster.

Poor partition field choices. Partitions can't be changed later, and queries that don't filter on them scan much more archived data. List your real archive queries before choosing.

Forgetting to switch connection strings. Your app's normal URI never sees archived data. If a feature suddenly "loses" old records after archiving, it's probably using the cluster-only string.

Using archive where TTL would do. If nobody ever needs the old data, archiving it just moves the cost. Delete it with a TTL index instead.

Relying on archive queries for latency-sensitive features. Archive reads go through object storage. Keep anything that must be fast in the cluster.

Conclusion

Atlas Online Archive lets you keep a small, hot working set in your cluster while moving older data to much cheaper storage, without building an export pipeline or giving up the ability to query history. A date-based rule, an index on that date, and a thoughtful choice of partition fields are most of the work. The payoff is smaller disks, smaller backups, and often a smaller cluster tier.

Pick your largest time-ordered collection, run the eligibility check above with a realistic cutoff, and write down the three queries your team actually runs against old data. If most of the collection is eligible and those queries filter on a couple of common fields, you have everything you need to configure your first archive.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Search: Adding Full-Text Search to Your App Without Elasticsearch

Atlas Search: Adding Full-Text Search to Your App Without Elasticsearch

The traditional way to add good search to a MongoDB app goes like this: stand up an Elasticsearch or OpenSearch cluster, write a sync process that co

Continue Reading