
Preventing NoSQL Injection Attacks in MongoDB Applications
A lot of developers pick up MongoDB and quietly assume injection is a SQL problem. There's no query string to concatenate, the driver takes objects, so what could go wrong? Quite a lot, it turns out. MongoDB queries are built from data structures, and if an attacker can control the shape of that structure instead of just the values inside it, they can rewrite your query just as effectively as a classic ' OR 1=1 -- payload.
NoSQL injection in MongoDB usually looks like an attacker sneaking a query operator such as $ne, $gt, or $regex into a place where your code expected a plain string. The result can be an authentication bypass, data leakage across tenants, or a query that pins your CPU for minutes. The good news is that the defenses are straightforward and mostly come down to one habit: never let untrusted input decide what kind of value it is.
This guide covers how operator injection works, the server-side JavaScript risks ($where, $function, $accumulator), injection through aggregation pipelines and sort or projection parameters, and the concrete defenses you should build into every Node.js and Python MongoDB app.
How Operator Injection Works
Here's a login route that looks reasonable at first glance:
app.post("/login", async (req, res) => {
const { username, password } = req.body;
const user = await db.collection("users").findOne({ username, password });
if (!user) return res.status(401).json({ error: "Invalid credentials" });
res.json({ ok: true, userId: user._id });
});
(Storing plaintext passwords is its own disaster, but bear with the example.) With express.json() enabled, an attacker doesn't have to send strings. They can send this body:
{
"username": "admin",
"password": { "$ne": null }
}
Your code builds the filter { username: "admin", password: { $ne: null } }, which matches the admin user as long as the password field exists. The attacker is in without knowing the password.
The same trick works through query strings. Express's default query parser (and many others) turns ?username=admin&password[$ne]=x into a nested object. So a GET endpoint that passes req.query fields straight into a filter is just as vulnerable.
It's Not Only Authentication
Operator injection shows up anywhere user input becomes part of a filter:
- Data exposure.
GET /api/orders?customerId[$gt]=turns "show me my orders" into "show me every order with a customerId". - Tenant escapes. If a tenant ID comes from the request rather than the authenticated session,
{ $exists: true }matches every tenant. - Denial of service.
{ "$regex": "^(a+)+$" }against a large unindexed field can burn CPU. MongoDB's regex engine has safeguards, but a slow regex on a big collection still hurts. - Blind extraction. With
$regexan attacker can probe a secret one character at a time: does the reset token start witha? Withab? The response (or its timing) leaks the answer.
Defense 1: Enforce Types at the Boundary
The single most effective defense is to make sure every value that ends up in a filter has the type you expect. If a field should be a string, reject anything that isn't a string before it reaches the database.
A minimal, explicit check:
app.post("/login", async (req, res) => {
const { username, password } = req.body;
if (typeof username !== "string" || typeof password !== "string") {
return res.status(400).json({ error: "Invalid input" });
}
const user = await db.collection("users").findOne({ username });
if (!user || !(await bcrypt.compare(password, user.passwordHash))) {
return res.status(401).json({ error: "Invalid credentials" });
}
res.json({ ok: true, userId: user._id });
});
Notice two changes. The type check blocks object payloads, and the password is no longer part of the query at all. Looking up by username and comparing a hash in application code removes the attack surface entirely.
Hand-written checks don't scale well, so in practice you want a schema validation library. Here's the same idea with Zod:
import { z } from "zod";
const LoginBody = z.object({
username: z.string().min(1).max(64),
password: z.string().min(8).max(256),
});
app.post("/login", async (req, res) => {
const parsed = LoginBody.safeParse(req.body);
if (!parsed.success) return res.status(400).json({ error: "Invalid input" });
const { username, password } = parsed.data;
// ...
});
Zod (or Joi, Yup, Valibot, or Pydantic in Python) rejects { "$ne": null } because it isn't a string. Validation libraries also strip unknown keys by default, which matters for the next section.
Casting ObjectIds
Many filters use _id. Converting the input with new ObjectId() is itself a type check, since it throws on anything that isn't a valid 24-character hex string or 12-byte value:
import { ObjectId } from "mongodb";
function parseId(value) {
if (typeof value !== "string" || !ObjectId.isValid(value)) return null;
return new ObjectId(value);
}
const id = parseId(req.params.id);
if (!id) return res.status(400).json({ error: "Bad id" });
const doc = await db.collection("posts").findOne({ _id: id });
ObjectId.isValid accepts some 12-character strings, which is why the typeof check and the explicit construction are both there. If you want to be strict, also check the string against /^[0-9a-f]{24}$/i.
Defense 2: Never Spread Request Objects Into Queries or Updates
A surprisingly common pattern is passing a whole request object into a query:
// Dangerous: the client controls every key and operator
const products = await db.collection("products").find(req.query).toArray();
Or into an update:
// Dangerous: the client can set any field, including role or tenantId
await db.collection("users").updateOne({ _id: userId }, { $set: req.body });
The update case is a mass assignment vulnerability. A user editing their display name can also send { "role": "admin" }. Build filters and updates field by field from validated input:
const ProfileUpdate = z
.object({
displayName: z.string().max(80).optional(),
bio: z.string().max(500).optional(),
})
.strict();
const data = ProfileUpdate.parse(req.body);
await db.collection("users").updateOne({ _id: userId }, { $set: data });
The .strict() call makes Zod reject unknown keys instead of silently stripping them, which is useful for spotting probing attempts in your logs.
Defense 3: Sanitize Keys as a Safety Net
Type validation is the primary defense. Stripping keys that start with $ or contain . is a useful second layer for apps where validation coverage is patchy, such as a large legacy Express codebase.
The express-mongo-sanitize middleware has been the popular choice for years, though it has compatibility issues with Express 5 because req.query became a getter. A small helper you own is often simpler:
function stripOperators(value) {
if (Array.isArray(value)) return value.map(stripOperators);
if (value && typeof value === "object" && !(value instanceof Date)) {
const clean = {};
for (const [key, val] of Object.entries(value)) {
if (key.startsWith("$") || key.includes(".")) continue;
clean[key] = stripOperators(val);
}
return clean;
}
return value;
}
Treat this as a backstop, not a strategy. Sanitizing turns { "$ne": null } into {}, which avoids the operator but still hands your query an object where it expected a string. Only type checking makes the failure explicit.
Mongoose's sanitizeFilter
If you use Mongoose, turn on sanitizeFilter. It wraps any nested object containing a key starting with $ in $eq, so the injected operator becomes a literal value to compare against:
import mongoose from "mongoose";
mongoose.set("sanitizeFilter", true);
// { password: { $ne: null } } becomes { password: { $eq: { $ne: null } } }
await User.findOne({ username, password: req.body.password });
When you genuinely need an operator in a sanitized query, wrap it with mongoose.trusted():
await User.find({ age: mongoose.trusted({ $gte: minAge }) });
Mongoose's schema casting also helps: a String path will try to cast its input, and strictQuery controls whether filters on fields not in the schema are dropped. Casting alone doesn't block objects with operators, though, which is why sanitizeFilter exists.
Server-Side JavaScript: $where, $function, and $accumulator
MongoDB can execute JavaScript on the server through the $where query operator and the $function and $accumulator aggregation operators. If any user input is concatenated into that JavaScript, you've created a code injection vulnerability inside your database:
// Never do this
db.collection("users").find({ $where: `this.name == '${req.query.name}'` });
A payload like ' || true || ' matches everything, and a payload like '; while(true){} ' ties up a JavaScript execution thread.
The fix is to avoid server-side JavaScript entirely. Almost everything $where does can be expressed with $expr and native aggregation operators, which are also much faster:
// Instead of $where: "this.spent > this.budget"
db.campaigns.find({ $expr: { $gt: ["$spent", "$budget"] } });
Then disable it at the server level so nobody can reintroduce it. In mongod.conf:
security:
javascriptEnabled: false
With that set, $where, $function, $accumulator, and mapReduce with JavaScript functions all fail. On Atlas, you can turn off server-side JavaScript in the cluster's additional settings. Note that server-side JavaScript is deprecated for $where, $function, and $accumulator as of MongoDB 8.0, which is another reason to move away from it now.
Injection Through Aggregation, Sort, and Projection
Filters aren't the only attack surface. Anything that turns user input into a pipeline stage or a field path is risky.
User-Controlled Field Names
A reporting endpoint that lets users pick a grouping field might build { $group: { _id: "$" + req.query.field } }. That's fine if field is restricted to known values, and a problem if the user can reference passwordHash or inject an expression object. Always map user choices to an allowlist:
const GROUP_FIELDS = {
country: "$address.country",
plan: "$plan",
status: "$status",
};
const groupBy = GROUP_FIELDS[req.query.groupBy];
if (!groupBy) return res.status(400).json({ error: "Unsupported groupBy" });
const rows = await db
.collection("customers")
.aggregate([{ $group: { _id: groupBy, count: { $sum: 1 } } }])
.toArray();
Sorting and Projection
Sort parameters like ?sort=createdAt are harmless until someone sends ?sort=passwordResetToken and learns the token ordering, or a projection parameter that includes sensitive fields. Apply the same allowlist approach:
const SORTS = {
newest: { createdAt: -1 },
oldest: { createdAt: 1 },
price: { price: 1 },
};
const sort = SORTS[req.query.sort] ?? SORTS.newest;
Regex Search Boxes
If you use user input in a $regex, escape it so users search for literal text rather than supplying patterns:
const escapeRegex = (s) => s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
const term = typeof req.query.q === "string" ? req.query.q.slice(0, 100) : "";
const results = await db
.collection("products")
.find({ name: { $regex: "^" + escapeRegex(term), $options: "i" } })
.limit(20)
.toArray();
Anchoring with ^ also lets MongoDB use an index for case-sensitive prefix searches. For anything beyond simple prefix matching, a text index or Atlas Search is a better tool; see regular expression queries in MongoDB for the performance side.
The Same Problems in Python
PyMongo isn't immune. Flask's request.get_json() and FastAPI's raw Request.json() return dicts that can contain $ keys. FastAPI with typed Pydantic models is safe by default, because a str field rejects a dict:
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from pymongo import AsyncMongoClient
app = FastAPI()
client = AsyncMongoClient("mongodb://localhost:27017")
users = client.app.users
class Login(BaseModel):
username: str = Field(min_length=1, max_length=64)
password: str = Field(min_length=8, max_length=256)
@app.post("/login")
async def login(body: Login):
user = await users.find_one({"username": body.username})
if not user or not verify_hash(body.password, user["password_hash"]):
raise HTTPException(status_code=401, detail="Invalid credentials")
return {"ok": True}
Sending {"username": "admin", "password": {"$ne": null}} to this endpoint returns a 422 validation error before any query runs. The danger comes back the moment you accept dict or Any types and pass them into find().
Limit the Blast Radius
Even with perfect input handling, assume something will slip through eventually and design so it doesn't matter much.
- Least-privilege database users. Your web app's user should have
readWriteon its own database, notdbAdminorroot. Create separate users for migrations and reporting. A read-only reporting service shouldn't be able to drop collections. - Scope every query from the session. Tenant IDs, user IDs, and roles come from the authenticated session or token, never from the request body or query string.
- Set time limits.
maxTimeMSon queries driven by user input caps the damage from expensive filters or regexes. - Validate at the database too. Schema validation with JSON Schema stops malformed documents (a
roleof{}, for example) from being written, even if application checks miss them. - Log and alert on rejected input. A spike in 400 responses with
$characters in the payload is a strong signal that someone is probing.
const results = await db
.collection("products")
.find(filter, { maxTimeMS: 2000 })
.limit(50)
.toArray();
Common Mistakes
Trusting the query string to contain strings. Most frameworks' query parsers support nested syntax like field[$ne]=x. If you don't check types, you'll get objects. Express 5 uses the simpler node:querystring-style parser by default, which reduces this risk, but Express 4 apps and many other frameworks still parse nested keys.
Relying only on sanitization middleware. Stripping $ keys is a good backup, but it leaves objects where strings should be and can break legitimate data such as keys in user-supplied JSON. Validate types first.
Putting secrets in the filter. Querying { username, password } or { token: req.query.token } makes the secret itself injectable. Look up by an identifier, then compare secrets in code using a constant-time comparison.
Building $where or $function strings from input. Even with escaping, it's code injection waiting to happen. Rewrite with $expr and disable server-side JavaScript.
Letting clients pick arbitrary sort, projection, or group fields. Map every user choice to an allowlist of known-safe values.
Conclusion
MongoDB doesn't have SQL, but it has something with the same failure mode: queries built from structures that attackers can shape if you let them. Operator injection through $ne and $regex, mass assignment through $set: req.body, and code injection through $where all come from treating request data as if it were already trusted. Enforce types with a schema library, build filters and updates field by field, use allowlists for field names, disable server-side JavaScript, and run your app with a least-privilege database user.
Pick one route in your app today that passes req.body or req.query into a MongoDB call, send it {"$ne": null} for a string field, and see what happens. If the request isn't rejected with a 400, add a schema to that route and then work outward from there.


