Type something to search...
Implementing Soft Deletes in MongoDB

Implementing Soft Deletes in MongoDB

A customer deletes a project, then emails support an hour later asking for it back. An admin removes a user account, and the finance team discovers their invoices now reference a document that no longer exists. A bug in a cleanup job wipes out a thousand records before anyone notices. Every one of these situations is easier to recover from if "delete" didn't actually delete anything.

That's the idea behind soft deletes: instead of removing a document, you mark it as deleted and filter it out of normal reads. The data stays around for restores, audits, and referential sanity, and a separate process purges it permanently when it's no longer needed. The concept takes one sentence. Doing it well takes some care, because the "deleted" filter needs to show up in every query, every index, and every uniqueness rule, and forgetting it once leaks deleted data back into your app.

This guide covers how to model a soft delete, how to index it with partial indexes, how to keep unique constraints working, how to apply the filter automatically in the Node.js driver and Mongoose, how to purge old deletions with a TTL index, and when a separate archive collection is the better design.

Modeling the Deleted State

The minimum is a flag. In practice you'll want to know when and by whom something was deleted, too:

{
  _id: ObjectId("66ee1a2b9c1d4e5f60718293"),
  email: "alice@example.com",
  name: "Alice Moreau",
  deleted: false,
  deletedAt: null,
  deletedBy: null
}

When the document is deleted:

db.users.updateOne(
  { _id: ObjectId("66ee1a2b9c1d4e5f60718293"), deleted: false },
  {
    $set: {
      deleted: true,
      deletedAt: new Date(),
      deletedBy: ObjectId("66e0f1a2b3c4d5e6f7081920"),
    },
  },
);

Including deleted: false in the filter makes the operation idempotent and preserves the original deletion time if someone clicks "delete" twice.

Why a Boolean and Not Just deletedAt?

A common design uses only deletedAt, where null means active. It's tidy, and a query like { deletedAt: null } matches documents where the field is either null or missing, which is convenient for older documents written before the field existed.

The problem shows up with indexes. The most useful soft-delete indexes are partial indexes that only cover active documents, and partial filter expressions support equality on a value but not "field is null or missing" in a way the planner can match reliably. A boolean that's always present gives you a simple equality (deleted: false) that works cleanly in both queries and partial indexes. Keep deletedAt for auditing and for the TTL purge covered later.

If you have existing documents without the flag, backfill it once:

db.users.updateMany(
  { deleted: { $exists: false } },
  { $set: { deleted: false, deletedAt: null, deletedBy: null } },
);

And make new documents always include it, either in your insert code or with a schema default.

Indexing for Soft Deletes

Once most queries include deleted: false, your indexes should account for it. You have two options.

Add deleted to compound indexes. An index on { deleted: 1, createdAt: -1 } supports "active users, newest first." This works, but the index still contains every deleted document, which wastes space and memory if deletions accumulate.

Use partial indexes. A partial index only includes documents that match a filter expression:

db.users.createIndex(
  { createdAt: -1 },
  { partialFilterExpression: { deleted: false }, name: "active_by_created" },
);

The index contains only active users, so it's smaller and faster to maintain. The catch: MongoDB only uses a partial index when the query's filter guarantees it's a subset of the partial filter. This query can use the index:

db.users.find({ deleted: false }).sort({ createdAt: -1 }).limit(20);

This one can't, because it might need deleted documents the index doesn't contain:

db.users.find({}).sort({ createdAt: -1 }).limit(20);

That's actually a nice property. It means forgetting the soft-delete filter in a query tends to show up as a slow query in your monitoring rather than silently returning deleted data from a fast index. Check your important queries with explain() to make sure they select the partial index.

Keeping Unique Constraints Working

Here's the problem that catches almost everyone. You have a unique index on email. Alice deletes her account (soft delete). A month later she signs up again with the same email, and the insert fails with a duplicate key error, because her old, deleted document still holds that email.

A partial unique index solves it by enforcing uniqueness only among active documents:

db.users.dropIndex("email_1");

db.users.createIndex(
  { email: 1 },
  {
    unique: true,
    partialFilterExpression: { deleted: false },
    name: "email_unique_active",
  },
);

Now there can be only one active user per email, and any number of deleted ones. Alice's new signup succeeds.

This has a consequence for restores. If Alice's new account is active and someone tries to restore the old one, the restore fails with a duplicate key error, which is the correct outcome: you can't have two active accounts with one email. Your restore flow should catch that error and explain the conflict instead of crashing. The post on handling duplicate key errors and unique constraints shows how to detect and report it.

If you need case-insensitive uniqueness too, add a collation to the same index. Partial filters and collations combine without issue.

Applying the Filter Automatically

The biggest risk with soft deletes is a query that forgets the filter. Relying on every developer to remember deleted: false in every query won't hold up. Centralize it.

A Repository Layer with the Node.js Driver

With the native driver, wrap the collection in a small module that adds the filter by default:

const ACTIVE = { deleted: false };

export function usersRepo(db) {
  const users = db.collection("users");

  return {
    findOne(filter, options) {
      return users.findOne({ ...filter, ...ACTIVE }, options);
    },
    find(filter, options) {
      return users.find({ ...filter, ...ACTIVE }, options);
    },
    countDocuments(filter = {}) {
      return users.countDocuments({ ...filter, ...ACTIVE });
    },
    updateOne(filter, update, options) {
      return users.updateOne({ ...filter, ...ACTIVE }, update, options);
    },

    async softDelete(id, actorId) {
      const res = await users.updateOne(
        { _id: id, ...ACTIVE },
        { $set: { deleted: true, deletedAt: new Date(), deletedBy: actorId } },
      );
      return res.modifiedCount === 1;
    },

    async restore(id) {
      const res = await users.updateOne(
        { _id: id, deleted: true },
        { $set: { deleted: false, deletedAt: null, deletedBy: null } },
      );
      return res.modifiedCount === 1;
    },

    // Explicit escape hatch for admin tools and audits
    findIncludingDeleted(filter, options) {
      return users.find(filter, options);
    },
  };
}

Spreading ACTIVE last means callers can't accidentally override it with their own deleted value. When you do need deleted documents, the escape hatch has a name that makes the intent obvious in code review.

Query Middleware in Mongoose

Mongoose lets you inject the filter with query middleware. The trick is to respect queries that explicitly ask about deleted, so that admin screens and restores still work:

import mongoose from "mongoose";

function softDeletePlugin(schema) {
  schema.add({
    deleted: { type: Boolean, default: false, index: true },
    deletedAt: { type: Date, default: null },
    deletedBy: { type: mongoose.Schema.Types.ObjectId, default: null },
  });

  function excludeDeleted() {
    if (this.getFilter().deleted === undefined) {
      this.where({ deleted: false });
    }
  }

  schema.pre(/^find/, excludeDeleted);
  schema.pre(["countDocuments", "updateOne", "updateMany"], excludeDeleted);

  schema.pre("aggregate", function () {
    const first = this.pipeline()[0];
    const mustBeFirst = first && ("$geoNear" in first || "$search" in first);
    this.pipeline().splice(mustBeFirst ? 1 : 0, 0, {
      $match: { deleted: false },
    });
  });

  schema.methods.softDelete = function (actorId) {
    this.deleted = true;
    this.deletedAt = new Date();
    this.deletedBy = actorId ?? null;
    return this.save();
  };

  schema.statics.restore = function (filter) {
    return this.updateMany(
      { ...filter, deleted: true },
      { $set: { deleted: false, deletedAt: null, deletedBy: null } },
    );
  };
}

const userSchema = new mongoose.Schema({ email: String, name: String });
userSchema.plugin(softDeletePlugin);
userSchema.index(
  { email: 1 },
  { unique: true, partialFilterExpression: { deleted: false } },
);

export const User = mongoose.model("User", userSchema);

Usage looks like normal Mongoose code:

await User.find({ name: /^Al/ }); // active only
await User.find({ deleted: true }); // explicitly deleted
await User.restore({ _id: userId }); // bypasses the filter because it sets deleted itself

The /^find/ pattern covers find, findOne, findOneAndUpdate, and the other findOneAnd* methods. The aggregate hook handles stages like $geoNear and $search, which must be first in a pipeline. Test the plugin with a few deliberate queries, because middleware coverage differs between methods and Mongoose versions.

Read-Only Access Through a View

For reporting tools, BI connectors, or teams that query the database directly, a view is a clean way to hide deleted data without trusting every consumer to filter:

db.createView("active_users", "users", [{ $match: { deleted: false } }]);

db.active_users.find({ email: "alice@example.com" });

Views are read-only and run their pipeline on every query, and queries against them can still use the underlying collection's indexes when the $match comes first. The post on views and on-demand materialized views covers the trade-offs.

Purging Deleted Documents with a TTL Index

Soft-deleted data shouldn't live forever. Storage costs money, indexes grow, and privacy rules may require real deletion after a period. A TTL index on deletedAt purges documents automatically:

db.users.createIndex(
  { deletedAt: 1 },
  { expireAfterSeconds: 60 * 60 * 24 * 30, name: "purge_deleted_after_30d" },
);

A TTL index only expires documents where the indexed field holds a date. Active users have deletedAt: null, so they're never touched. Once a document is soft-deleted, the background TTL monitor removes it roughly 30 days after deletedAt. The monitor runs periodically (about once a minute), so expiry isn't instant, but it's reliable.

If restoring sets deletedAt back to null, the TTL clock resets too. That's exactly the behavior you want. See TTL indexes for expiring old data for more on how the monitor works.

Handling Related Documents

Soft deletes raise a design question that hard deletes hide: what happens to documents that reference the deleted one?

  • Leave references alone when the related data has independent meaning. Invoices should keep pointing at the customer who placed the order, deleted or not. Your UI can display "Deleted user" when the lookup finds a soft-deleted document.
  • Cascade the soft delete when children make no sense without the parent. Deleting a project might soft-delete its tasks in the same operation. Use a transaction if both updates must succeed together.
  • Filter on join when using $lookup. A plain localField/foreignField lookup returns deleted documents too. Use a pipeline form to exclude them:
db.orders.aggregate([
  {
    $lookup: {
      from: "users",
      localField: "userId",
      foreignField: "_id",
      pipeline: [
        { $match: { deleted: false } },
        { $project: { name: 1, email: 1 } },
      ],
      as: "user",
    },
  },
]);

Decide these rules per relationship and write them down. Inconsistent cascade behavior is where soft-delete systems get confusing.

The Alternative: An Archive Collection

Instead of flagging documents in place, you can move them to a separate collection. The main collection stays clean, and nothing needs a deleted filter:

export async function archiveUser(client, userId, actorId) {
  const db = client.db("app");

  await client.withSession((session) =>
    session.withTransaction(async () => {
      const user = await db
        .collection("users")
        .findOneAndDelete({ _id: userId }, { session });
      if (!user) return;

      await db
        .collection("users_archive")
        .insertOne(
          { ...user, archivedAt: new Date(), archivedBy: actorId },
          { session },
        );
    }),
  );
}

The transaction guarantees the document is never in both collections or neither. Here's how the two approaches compare:

ConcernFlag in placeArchive collection
Every query needs a filterYesNo
Unique indexesNeed partial indexesWork as normal
RestoreOne updateMove back (transaction)
References to deleted documentsStill resolveBreak unless you check the archive
Delete costOne updateDelete plus insert in a transaction
Index size on the main collectionIncludes deleted docs unless partialOnly active docs

The archive approach is a good fit when deletions are rare, restores are rarer, and you care about keeping the hot collection lean. The flag approach is better when restores are common, references must keep resolving, or you want the simplest write path.

Soft Deletes Are Not Erasure

A soft delete is a product feature, not a privacy guarantee. If a user exercises a legal right to erasure, keeping their personal data with deleted: true doesn't satisfy it. Plan for a separate hard-delete or anonymization path that removes or scrubs personal fields, and remember that data also lives in backups, analytics copies, and logs.

Also keep in mind that downstream consumers see soft deletes as updates. If you use change streams to sync data elsewhere, a soft delete arrives as an update event with deleted: true in the updated fields, not as a delete event. Handle that explicitly in your consumers.

Common Pitfalls

Forgetting the filter in one query. One missed deleted: false in a search endpoint and deleted records reappear. Centralize the filter in a repository or middleware, and grep for direct collection access in code review.

Keeping a plain unique index. Users can't reuse an email or username after deletion. Replace global unique indexes with partial unique indexes scoped to active documents.

Partial indexes that never get used. If a query doesn't include deleted: false exactly, the planner can't choose the partial index. Confirm with explain().

Using deletedAt: null for everything. It's convenient for queries, but it doesn't combine well with partial indexes. Use an explicit boolean for filtering and deletedAt for auditing and TTL.

Keeping deleted data forever. Without a purge, deleted documents quietly dominate your storage. Add a TTL index on deletedAt or a scheduled cleanup job from day one.

Treating soft delete as compliance. Real erasure requests need a hard delete or anonymization path.

Conclusion

Soft deletes give you undo, audit trails, and protection against accidental data loss for the price of one field and a filter. The implementation details are what separate a solid system from a leaky one: an always-present boolean, partial indexes that cover only active documents, partial unique indexes so identifiers can be reused, a single place in your code that applies the filter, and a TTL index that eventually removes what nobody needs.

Pick one collection where users can delete things, add the deleted, deletedAt, and deletedBy fields with a backfill, and replace its unique index with a partial one. Then route every read through a repository or plugin before you add soft deletes anywhere else.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading