
Text Search in MongoDB: Creating and Querying Text Indexes
Your app needs a search box. The first version usually looks like find({ title: { $regex: userInput, $options: "i" } }), and for a few hundred documents it seems fine. Then the collection grows, the query starts scanning every document, and users notice that searching "run" doesn't find "running," that "the" matches everything, and that searching two words only works if they appear side by side in exactly that order.
A text index fixes most of that without adding any new infrastructure. It breaks string fields into words, drops common stop words, reduces words to their stems, and lets you query with the $text operator. Results can be ranked by relevance, and multiple fields can be weighted differently. It runs inside MongoDB itself, on every edition and deployment type.
This guide covers creating text indexes, the $text query syntax (terms, phrases, and negation), relevance scoring and field weights, languages, using text search in aggregations and application code, and the limitations that tell you when it's time to move to Atlas Search.
Creating a Text Index
A text index is created by giving one or more fields the index type "text":
db.articles.createIndex({ title: "text", body: "text" });
Let's add some data to work with:
db.articles.insertMany([
{
title: "Running a Coffee Shop",
body: "Lessons learned from five years behind the espresso machine.",
category: "business",
},
{
title: "Espresso at Home",
body: "How to pull a great shot without a commercial machine.",
category: "coffee",
},
{
title: "Marathon Training Plan",
body: "A sixteen week plan for first-time runners.",
category: "fitness",
},
{
title: "Decaf Myths",
body: "Why decaf coffee deserves a second look.",
category: "coffee",
},
]);
When MongoDB indexes these documents, it tokenizes each field into words, lowercases them, removes stop words like "a," "the," and "from," and stems what's left. "Running" and "runners" both stem to "run," which is why a search for "run" matches the first and third articles.
There's one hard rule: a collection can have only one text index. That single index can cover as many fields as you like, but you can't create a second one. Plan it to include every field you want searchable.
Indexing Every String Field
If your documents have unpredictable shapes, you can use a wildcard text index that indexes every field containing a string:
db.notes.createIndex({ "$**": "text" });
This is convenient for prototypes but makes the index large and gives you less control over relevance. For production collections, list the fields explicitly.
Querying With $text
Text queries use the $text operator with a $search string:
db.articles.find({ $text: { $search: "coffee" } }, { title: 1, _id: 0 });
[{ title: "Decaf Myths" }, { title: "Running a Coffee Shop" }];
Notice what didn't match: "Espresso at Home" is obviously about coffee, but the word never appears. A text index matches words, not concepts. (Synonyms are one of the things Atlas Search adds.)
Multiple Terms Mean OR
When the search string contains multiple words, MongoDB matches documents containing any of them:
db.articles.find(
{ $text: { $search: "espresso marathon" } },
{ title: 1, _id: 0 },
);
[
{ title: "Marathon Training Plan" },
{ title: "Espresso at Home" },
{ title: "Running a Coffee Shop" },
];
This surprises people who expect search to narrow results as they type more words. It widens them instead. Relevance scoring (covered below) pushes documents matching more terms to the top, which is usually what you want for a search box, but if you need every word to be present, use phrases.
Phrases
Wrap words in escaped double quotes to require an exact phrase:
db.articles.find({ $text: { $search: '"coffee shop"' } }, { title: 1, _id: 0 });
[{ title: "Running a Coffee Shop" }];
If the search string contains a phrase and other terms, documents must contain the phrase. Multiple phrases are combined with AND in recent versions, so "\"espresso machine\" \"five years\"" requires both. Phrase matching is done against the original text after the index narrows the candidates, so it's more expensive than a plain term query.
To require several individual words (a logical AND), quote each one:
db.articles.find({ $text: { $search: '"coffee" "decaf"' } });
Negation
Prefix a term with a hyphen to exclude documents containing it:
db.articles.find({ $text: { $search: "coffee -decaf" } }, { title: 1, _id: 0 });
[{ title: "Running a Coffee Shop" }];
A search made up only of negated terms matches nothing, since there's no positive term to find candidates with.
Combining $text With Other Filters
$text combines with ordinary filters in the same query:
db.articles.find({
$text: { $search: "coffee espresso" },
category: "coffee",
});
MongoDB uses the text index to find candidates and then applies the category filter. If you always filter on a particular equality field, you can put it in the text index as a prefix, which lets the index do the filtering:
db.products.createIndex({ storeId: 1, name: "text", description: "text" });
db.products.find({ storeId: 42, $text: { $search: "lamp" } });
With a prefix field in the index, every $text query on that collection must include an equality match on storeId. That's a good fit for multi-tenant apps where every search is scoped to one tenant, and a bad fit if you sometimes search across all tenants.
Relevance Scoring
Every document matched by $text gets a relevance score based on how often the terms appear and in which fields. To see it and sort by it, use $meta: "textScore":
db.articles
.find(
{ $text: { $search: "coffee espresso machine" } },
{ title: 1, _id: 0, score: { $meta: "textScore" } },
)
.sort({ score: { $meta: "textScore" } });
[
{ title: "Espresso at Home", score: 1.625 },
{ title: "Running a Coffee Shop", score: 1.4 },
{ title: "Decaf Myths", score: 0.6666666666666666 },
];
Without an explicit sort, results come back in no particular order. Always sort by textScore when showing results to users. In recent versions you can sort by the score without also projecting it, but including it in the projection is handy for debugging.
The scores above are illustrative; yours will differ slightly depending on field lengths and index options. The exact numbers aren't meaningful on their own and can change as you edit documents. Use them for ordering, not as a percentage or threshold shown to users.
Field Weights
By default every indexed field has a weight of 1. A match in the title usually matters more than a match deep in the body, so weight it higher when creating the index:
db.articles.dropIndex("title_text_body_text");
db.articles.createIndex(
{ title: "text", tags: "text", body: "text" },
{
weights: { title: 10, tags: 5, body: 1 },
name: "ArticleSearch",
},
);
A term in the title now contributes ten times as much to the score as the same term in the body. Weights are fixed at index creation, so changing them means dropping and rebuilding the text index. Giving the index an explicit name makes that easier later, since generated names for text indexes get long.
Languages, Stemming, and Stop Words
Stemming and stop words depend on language. The default is English. You can set a different default for the whole index:
db.recipes.createIndex(
{ title: "text", instructions: "text" },
{ default_language: "spanish" },
);
MongoDB supports a list of languages for stemming (including English, French, German, Spanish, Portuguese, Italian, Dutch, Russian, and several Scandinavian languages). Languages like Chinese, Japanese, and Korean aren't supported by the built-in tokenizer, which splits on whitespace and punctuation.
If a collection mixes languages, each document can declare its own with a language field:
db.recipes.insertMany([
{
title: "Tortilla de patatas",
instructions: "Cortar las patatas...",
language: "spanish",
},
{
title: "Shepherd's pie",
instructions: "Brown the lamb mince...",
language: "english",
},
]);
The field name is configurable with the language_override index option, which is important if your documents already use language for something else, like a programming language on a code snippet. If that field holds a value MongoDB doesn't recognize as a supported language, the insert fails, which is a confusing error the first time you see it:
db.snippets.createIndex(
{ title: "text", code: "text" },
{ language_override: "textLanguage" },
);
Setting default_language: "none" disables stemming and stop words entirely. Words are only tokenized and lowercased. That's useful for product codes, usernames, or text where stemming causes false matches.
At query time, you can specify the language for the search string:
db.recipes.find({ $text: { $search: "patatas", $language: "spanish" } });
Case and Diacritics
Text indexes are case-insensitive and, with the current index version (version 3), diacritic-insensitive by default. "Café" matches "cafe." You can opt into stricter matching per query:
db.articles.find({
$text: { $search: "Café", $caseSensitive: true, $diacriticSensitive: true },
});
These options make queries slower because MongoDB has to verify candidates against the original text, so use them only when you need them.
Text Search in Aggregation Pipelines
In an aggregation, $text goes inside a $match that must be the first stage of the pipeline. After that, you can use the score like any other field:
db.articles.aggregate([
{ $match: { $text: { $search: "coffee espresso" } } },
{ $set: { score: { $meta: "textScore" } } },
{ $match: { score: { $gt: 1 } } },
{ $sort: { score: -1 } },
{
$facet: {
results: [{ $limit: 10 }, { $project: { title: 1, score: 1 } }],
byCategory: [{ $group: { _id: "$category", count: { $sum: 1 } } }],
},
},
]);
This returns the top ten results and a count per category in one round trip, which is a nice building block for a search page with filters. The post on using $facet for multi-faceted search results builds this out further.
Text Search From Application Code
Here's a small Express endpoint using the Node.js driver:
import express from "express";
import { MongoClient } from "mongodb";
const client = new MongoClient(process.env.MONGODB_URI);
const articles = client.db("blog").collection("articles");
const app = express();
app.get("/search", async (req, res) => {
const q = String(req.query.q ?? "")
.trim()
.slice(0, 200);
if (!q) return res.json({ results: [] });
const page = Math.max(1, Number(req.query.page) || 1);
const pageSize = 10;
const results = await articles
.find(
{ $text: { $search: q } },
{ projection: { title: 1, category: 1, score: { $meta: "textScore" } } },
)
.sort({ score: { $meta: "textScore" } })
.skip((page - 1) * pageSize)
.limit(pageSize)
.toArray();
res.json({ page, results });
});
app.listen(3000);
The search string is passed straight into $search, which is safe because it's treated as text, not as a query object. Do make sure it's a string, though. If req.query.q could be an object, someone can inject operators, which is covered in preventing NoSQL injection. The String() conversion and length cap handle that here.
And the same query in PyMongo:
from pymongo import MongoClient
articles = MongoClient("mongodb://localhost:27017")["blog"]["articles"]
cursor = (
articles.find(
{"$text": {"$search": "coffee -decaf"}},
{"title": 1, "score": {"$meta": "textScore"}},
)
.sort([("score", {"$meta": "textScore"})])
.limit(10)
)
for doc in cursor:
print(f"{doc['score']:.2f} {doc['title']}")
Limitations to Know
Text indexes are useful, but they're a basic search engine. These limitations are the ones that most often push teams toward something more capable:
| Limitation | What it means in practice |
|---|---|
| One text index per collection | All searchable fields must be in one index with fixed weights |
| No partial word or prefix matching | "espr" doesn't match "espresso," so no search-as-you-type |
| No fuzzy matching | Typos like "expresso" return nothing |
| No synonyms | "couch" won't find "sofa" |
| Limited scoring control | Weights are fixed at creation; no boosting by popularity or recency in the index |
$text must be first in pipelines | Can't search the output of earlier stages |
One $text per query, no hint() | Can't be combined with index hints or used inside $nor |
| Not supported on views | Build the search on the underlying collection |
| Index size | Text indexes on long fields can be large and slow down writes |
A couple of query restrictions deserve extra detail. If you use $text inside an $or, every clause of the $or must be supported by an index. And text search can't be combined with $near in the same query.
When to Move to Atlas Search
If you need autocomplete, typo tolerance, synonyms, highlighting, or custom relevance, the built-in text index won't get you there. Atlas Search is a Lucene-based search engine integrated with MongoDB that supports all of those through the $search aggregation stage, and it keeps itself in sync with your collection automatically. Recent MongoDB releases have also started bringing the same search engine to self-managed Community and Enterprise deployments, so check the current docs for availability on your version. Atlas Search without Elasticsearch walks through the setup.
For internal tools, admin panels, modest catalogs, and anywhere "find documents containing these words" is good enough, the text index remains a perfectly good choice with zero extra infrastructure.
Common Pitfalls
Forgetting to sort by score. Unsorted $text results come back in index order, which looks random to users. Sort by { $meta: "textScore" }.
Expecting multiple words to narrow results. Terms are ORed. Quote terms or phrases when every word must appear.
Trying to add a second text index. The create fails. Drop the existing one and create a new index that covers all fields.
Colliding with a language field. If your documents use language for anything other than a supported human language, inserts fail. Set language_override to a different field name.
Using text search for autocomplete. Text indexes match whole stemmed words, not prefixes. Use Atlas Search's autocomplete, or a prefix-anchored regex on an indexed field for simple cases.
Indexing huge fields you rarely search. Every word in every indexed field becomes an index entry. Leave out long fields that don't help relevance.
Conclusion
Text indexes give you real word-based search inside MongoDB with a single createIndex call. You get stemming, stop words, relevance scores, phrase and negation syntax, per-field weights, and multi-language support, all queried through $text in find() or at the start of an aggregation. The trade-offs are one index per collection and no fuzzy matching, prefixes, or synonyms.
Replace one $regex search in your app with a text index today: create the index with weights that favor your title field, switch the query to $text sorted by textScore, and compare both the speed and the quality of the results. If users immediately ask for autocomplete, that's your signal to look at Atlas Search next.


