Type something to search...
Atlas Search: Adding Full-Text Search to Your App Without Elasticsearch

Atlas Search: Adding Full-Text Search to Your App Without Elasticsearch

The traditional way to add good search to a MongoDB app goes like this: stand up an Elasticsearch or OpenSearch cluster, write a sync process that copies every insert, update, and delete from MongoDB into it, handle the failures when that sync falls behind, and then query two different systems with two different query languages. It works, but you've doubled your operational surface area to power one search box.

Atlas Search removes most of that. It's a Lucene-based search engine built into MongoDB Atlas that indexes your collections directly and keeps itself in sync automatically. You query it with a $search stage in an ordinary aggregation pipeline, which means search results can flow straight into $project, $lookup, $facet, and everything else you already know. You get the features users expect from modern search, like autocomplete, fuzzy matching, synonyms, facets, and highlighting, without a second database.

This guide covers creating search indexes, writing $search queries, compound queries with filters and boosts, autocomplete, fuzzy matching, highlighting, facets, pagination, and the operational details that matter once it's in production.

How Atlas Search Works

When you create a search index, Atlas runs a process called mongot alongside your cluster's mongod nodes (or on dedicated Search Nodes, which we'll get to). mongot builds a Lucene index from your collection, then tails the change stream to apply every write. When your pipeline contains $search, mongod forwards that stage to mongot, gets back matching document IDs and scores, and continues the pipeline with the full documents.

Two consequences are worth knowing from the start:

  • Search indexes are eventually consistent. A document you just inserted typically shows up in search within a second or so, but not necessarily in the same request. Don't write a document and immediately assert it's searchable in a test without polling.
  • Search indexes are separate from regular indexes. They don't speed up find() queries, and regular indexes don't help $search. You manage them with their own commands.

Atlas Search is available on all Atlas cluster tiers, including the free M0 tier and Flex clusters, with limits on the number of search indexes on the smaller tiers. That makes it easy to prototype for free. For local development, the Atlas CLI can run a local Atlas deployment in Docker with search included:

atlas deployments setup local-dev --type local
atlas deployments connect local-dev --connectWith mongosh

There's also a mongodb/mongodb-atlas-local Docker image if you'd rather wire it into docker compose yourself.

Creating a Search Index

A search index is defined by a JSON document describing which fields to index and how. The simplest definition uses dynamic mapping, which indexes every field with default settings:

db.products.createSearchIndex("default", {
  mappings: { dynamic: true },
});

Dynamic mappings are great for exploring, but in production you usually want static mappings: you choose the fields, their types, and their analyzers. That keeps the index smaller and gives you control over matching. Here's a definition for a product catalog:

{
  "mappings": {
    "dynamic": false,
    "fields": {
      "name": [
        { "type": "string", "analyzer": "lucene.english" },
        {
          "type": "autocomplete",
          "tokenization": "edgeGram",
          "minGrams": 2,
          "maxGrams": 15
        }
      ],
      "description": { "type": "string", "analyzer": "lucene.english" },
      "brand": [{ "type": "string" }, { "type": "token" }],
      "category": { "type": "token" },
      "price": { "type": "number" },
      "rating": { "type": "number" },
      "inStock": { "type": "boolean" },
      "createdAt": { "type": "date" }
    }
  }
}

A few things to notice:

  • name is indexed twice, once as searchable text with English stemming and once for autocomplete. A field can have multiple mappings.
  • category uses the token type, which indexes the whole value as one term. That's what you want for exact filters, sorting, and facets on strings.
  • Numbers, booleans, and dates get their own types so you can filter and facet on them inside $search.

Create it from mongosh:

db.products.createSearchIndex("products", {
  mappings: {
    dynamic: false,
    fields: {
      name: [
        { type: "string", analyzer: "lucene.english" },
        {
          type: "autocomplete",
          tokenization: "edgeGram",
          minGrams: 2,
          maxGrams: 15,
        },
      ],
      description: { type: "string", analyzer: "lucene.english" },
      category: { type: "token" },
      price: { type: "number" },
      inStock: { type: "boolean" },
    },
  },
});

Or from Node.js, which is handy in a migration script so your index definitions live in version control:

await db.collection("products").createSearchIndex({
  name: "products",
  definition: {
    mappings: { dynamic: false, fields: {/* ...same as above... */} },
  },
});

You can check build status with db.products.getSearchIndexes(). The index is queryable once its status is READY. You can also create and edit indexes in the Atlas UI, which has a visual editor, or with atlas clusters search indexes create from the CLI.

Your First $search Query

$search must be the first stage of the pipeline. The text operator is the workhorse:

db.products.aggregate([
  {
    $search: {
      index: "products",
      text: { query: "wireless headphones", path: ["name", "description"] },
    },
  },
  { $limit: 5 },
  { $project: { name: 1, price: 1, score: { $meta: "searchScore" } } },
]);
[
  {
    _id: ObjectId("..."),
    name: "Wireless Noise-Cancelling Headphones",
    price: 199,
    score: 6.41,
  },
  {
    _id: ObjectId("..."),
    name: "Studio Headphones (Wired)",
    price: 129,
    score: 3.02,
  },
  { _id: ObjectId("..."), name: "Wireless Earbuds", price: 89, score: 2.87 },
];

Results come back sorted by relevance automatically, highest score first. The lucene.english analyzer handles stemming, so "headphone" matches "headphones," and multiple terms are scored so documents matching more of them rank higher.

Compound Queries: Must, Should, Filter, MustNot

Real search combines relevance with constraints. The compound operator has four clauses:

ClauseMust match?Affects score?Use for
mustYesYesThe main query terms
shouldNoYesBoosting preferred matches
filterYesNoHard constraints like price or stock
mustNotMust notNoExclusions

Here's a search for headphones under $200 that are in stock, preferring highly rated products and matches in the name:

db.products.aggregate([
  {
    $search: {
      index: "products",
      compound: {
        must: [
          { text: { query: "headphones", path: ["name", "description"] } },
        ],
        should: [
          {
            text: {
              query: "headphones",
              path: "name",
              score: { boost: { value: 3 } },
            },
          },
          { range: { path: "rating", gte: 4.5 } },
        ],
        filter: [
          { range: { path: "price", lte: 200 } },
          { equals: { path: "inStock", value: true } },
        ],
        mustNot: [{ equals: { path: "category", value: "refurbished" } }],
      },
    },
  },
  { $limit: 20 },
  {
    $project: { name: 1, price: 1, rating: 1, score: { $meta: "searchScore" } },
  },
]);

Putting price and stock in filter rather than a later $match matters for performance. Filtering inside $search happens in the Lucene index, so mongot only returns documents that qualify. A $match after $search has to fetch every search hit from the database first and then throw most of them away.

Fuzzy Matching for Typos

Users misspell things. Add fuzzy to a text operator to tolerate small errors:

{
  $search: {
    index: "products",
    text: {
      query: "hedphones",
      path: "name",
      fuzzy: { maxEdits: 1, prefixLength: 2 }
    }
  }
}

maxEdits is the number of single-character changes allowed (1 or 2). prefixLength requires the first N characters to match exactly, which cuts down on noise and speeds up the query considerably. With maxEdits: 1, "hedphones" still finds "headphones." Two edits catches more typos but also more false positives, so start with one.

A common pattern is to put an exact text match in should with a boost and a fuzzy match in must, so exact matches outrank fuzzy ones.

Autocomplete

The autocomplete operator powers search-as-you-type. It queries the autocomplete-typed mapping we defined on name:

db.products.aggregate([
  {
    $search: {
      index: "products",
      autocomplete: {
        query: "noi",
        path: "name",
        fuzzy: { maxEdits: 1, prefixLength: 1 },
      },
    },
  },
  { $limit: 8 },
  { $project: { _id: 0, name: 1 } },
]);
[
  { name: "Wireless Noise-Cancelling Headphones" },
  { name: "Noise Machine for Sleep" },
];

The edgeGram tokenization indexes prefixes of each word (from minGrams to maxGrams characters), so "noi" matches the start of "Noise." If you want matches in the middle of words too, nGram tokenization does that, at the cost of a much larger index. Keep maxGrams reasonable, around 15, since longer prefixes rarely help and grow the index quickly.

In a UI, debounce the input by 150 to 250 milliseconds and only send a query once the user has typed two or three characters.

Highlighting Matches

To show users why a result matched, request highlights and project them with $meta: "searchHighlights":

db.products.aggregate([
  {
    $search: {
      index: "products",
      text: { query: "battery life", path: "description" },
      highlight: { path: "description" },
    },
  },
  { $limit: 3 },
  { $project: { name: 1, highlights: { $meta: "searchHighlights" } } },
]);
{
  name: "Wireless Noise-Cancelling Headphones",
  highlights: [
    {
      path: "description",
      score: 1.38,
      texts: [
        { value: "Up to 30 hours of ", type: "text" },
        { value: "battery", type: "hit" },
        { value: " ", type: "text" },
        { value: "life", type: "hit" },
        { value: " on a single charge.", type: "text" }
      ]
    }
  ]
}

Each texts array alternates plain text and hits, so your front end can wrap hit segments in a mark element. Build that markup from the pieces rather than inserting raw HTML strings, and you avoid injecting anything unsafe from your data.

Facets With $searchMeta

A search page usually shows counts next to filters: "Headphones (42), Speakers (17)." $searchMeta returns metadata instead of documents, and its facet collector computes those counts over the whole result set:

db.products.aggregate([
  {
    $searchMeta: {
      index: "products",
      facet: {
        operator: {
          text: { query: "wireless", path: ["name", "description"] },
        },
        facets: {
          categories: { type: "string", path: "category", numBuckets: 10 },
          priceRanges: {
            type: "number",
            path: "price",
            boundaries: [0, 50, 100, 200, 500],
            default: "500+",
          },
        },
      },
    },
  },
]);
[
  {
    count: { lowerBound: Long("63") },
    facet: {
      categories: {
        buckets: [
          { _id: "headphones", count: Long("28") },
          { _id: "speakers", count: Long("19") },
          { _id: "accessories", count: Long("16") },
        ],
      },
      priceRanges: {
        buckets: [
          { _id: 0, count: Long("12") },
          { _id: 50, count: Long("21") },
          { _id: 100, count: Long("22") },
          { _id: 200, count: Long("8") },
        ],
      },
    },
  },
];

String facets require the field to be indexed as token, which is why category got that type. To get results and facets in a single query, run $search with the same facet definition and read the metadata from the $$SEARCH_META variable in a later stage, for example inside $facet. Doing it in one pipeline saves a round trip and guarantees the counts match the results.

Sorting and Pagination

By default results are sorted by relevance. You can sort by indexed fields inside $search too, which is much faster than a $sort stage afterward:

{
  $search: {
    index: "products",
    text: { query: "headphones", path: "name" },
    sort: { price: 1 }
  }
}

For pagination, $skip and $limit work, but deep pages get slower because the engine still has to find and skip every earlier result. For infinite scroll or "next page" buttons, use searchAfter with a pagination token:

const first = await products
  .aggregate([
    {
      $search: {
        index: "products",
        text: { query: "headphones", path: "name" },
        sort: { price: 1 },
      },
    },
    { $limit: 20 },
    {
      $project: { name: 1, price: 1, token: { $meta: "searchSequenceToken" } },
    },
  ])
  .toArray();

const lastToken = first.at(-1).token;

const next = await products
  .aggregate([
    {
      $search: {
        index: "products",
        text: { query: "headphones", path: "name" },
        sort: { price: 1 },
        searchAfter: lastToken,
      },
    },
    { $limit: 20 },
    {
      $project: { name: 1, price: 1, token: { $meta: "searchSequenceToken" } },
    },
  ])
  .toArray();

This is the same idea as range-based pagination for regular queries, which skip/limit vs. range-based pagination explains in depth.

Wiring It Into an API

Here's an Express endpoint that ties together a compound query, optional filters, and highlights:

import express from "express";
import { MongoClient } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
const products = client.db("shop").collection("products");
const app = express();

app.get("/api/search", async (req, res) => {
  const q = String(req.query.q ?? "")
    .trim()
    .slice(0, 100);
  const category = req.query.category ? String(req.query.category) : null;
  const maxPrice = Number(req.query.maxPrice) || null;
  if (q.length < 2) return res.json({ results: [] });

  const filter = [{ equals: { path: "inStock", value: true } }];
  if (category) filter.push({ equals: { path: "category", value: category } });
  if (maxPrice) filter.push({ range: { path: "price", lte: maxPrice } });

  const results = await products
    .aggregate([
      {
        $search: {
          index: "products",
          compound: {
            must: [
              {
                text: {
                  query: q,
                  path: ["name", "description"],
                  fuzzy: { maxEdits: 1, prefixLength: 2 },
                },
              },
            ],
            should: [
              {
                text: {
                  query: q,
                  path: "name",
                  score: { boost: { value: 3 } },
                },
              },
            ],
            filter,
          },
          highlight: { path: "description" },
        },
      },
      { $limit: 20 },
      {
        $project: {
          name: 1,
          price: 1,
          category: 1,
          score: { $meta: "searchScore" },
          highlights: { $meta: "searchHighlights" },
        },
      },
    ])
    .toArray();

  res.json({ results });
});

app.listen(3000);

Note that user input only ever goes into query strings and value fields, never into the structure of the stage. Coercing everything with String() and Number() keeps a crafted query string from injecting operators.

Synonyms

Atlas Search can treat words as equivalent using a synonym mapping backed by a regular collection:

db.search_synonyms.insertMany([
  { mappingType: "equivalent", synonyms: ["couch", "sofa", "settee"] },
  { mappingType: "explicit", input: ["tv"], synonyms: ["tv", "television"] },
]);

Reference that collection in the index definition under synonyms, then add synonyms: "<mapping name>" to a text operator. Because synonyms live in a collection, your merchandising team can update them without a deploy. Note that a text query can use either synonyms or fuzzy, not both at once, so combine them with compound if you need both.

Operational Considerations

Resource sharing. By default mongot runs on the same nodes as your database, competing for CPU and memory. On dedicated clusters (M10 and up) you can deploy Search Nodes, separate nodes that run only search, so heavy search traffic doesn't slow down your writes and search can scale independently.

Index size. Autocomplete and nGram mappings can make an index many times larger than the source data. Map only the fields you search, and check the index size in the Atlas UI after building.

Changing index definitions. Updating a search index triggers a rebuild in the background. The old index keeps serving queries until the new one is ready, so updates don't cause downtime, but a rebuild on a large collection uses significant resources.

Cost. There's no separate license, but search consumes cluster resources (or Search Node resources), so factor it into sizing. It's still usually far cheaper than running and syncing a separate search cluster.

Common Pitfalls

Using $match for filters after $search. It works but forces MongoDB to fetch every hit. Put constraints in compound.filter.

Leaving dynamic mapping on in production. Every string field gets indexed, including ones nobody searches. Switch to a static mapping.

Faceting on a string mapping. String facets need the token type. Add it as a second mapping on the field.

Querying a field the index doesn't include. With static mappings, $search on an unmapped path returns nothing, without an error. If results are empty, check the definition first.

Testing without waiting for sync. Search is eventually consistent. In integration tests, poll until the document is searchable instead of asserting immediately.

Deep pagination with $skip. Use searchAfter tokens for anything past the first few pages.

Conclusion

Atlas Search gives you production-grade search inside the database you already run. Define a static search index, query it with $search at the start of an aggregation pipeline, and use compound to combine relevance with fast in-index filters. Autocomplete, fuzzy matching, highlighting, synonyms, and facets are all configuration rather than infrastructure, and there's no sync process to babysit. If you're coming from basic $text queries, the guide to text indexes is a useful comparison point.

Start with one collection your users search most. Create a search index with dynamic mapping on a free M0 cluster or a local Atlas deployment, try a few $search queries in mongosh, then replace dynamic mapping with a static one covering just the fields that matter.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading