
Mongoose Schemas, Models, and Middleware: A Practical Guide
Mongoose is easy to start with and surprisingly deep once you're in. Defining a schema and calling Model.find() takes five minutes. Understanding why your pre('save') hook didn't run on updateOne, why a validator was skipped, or why this is undefined inside a middleware function takes a bit longer.
Those surprises all come from the same place: Mongoose has three layers that work together. Schemas describe the shape and rules of your documents. Models compile a schema into a class with a query API. Middleware hooks into specific operations on documents, queries, and aggregations. Once you know which layer owns what, Mongoose becomes predictable.
This guide walks through each layer with a realistic example: a small blog with users and posts. It covers schema types and options, validation, virtuals, instance and static methods, the different kinds of middleware and when each runs, plus the pitfalls that catch most developers. The examples target Mongoose 8.x.
Setting Up
npm install mongoose
// db.js
import mongoose from "mongoose";
export async function connect() {
await mongoose.connect(process.env.MONGODB_URI, { dbName: "blog" });
}
Mongoose buffers model operations until the connection is ready, so you can define models before connecting. In production you'll want to connect at startup and fail fast if it doesn't work.
Schemas: Describing Your Documents
A schema maps field names to types and rules. Here's a user schema that uses most of the common options:
// models/User.js
import mongoose from "mongoose";
const { Schema } = mongoose;
const userSchema = new Schema(
{
email: {
type: String,
required: [true, "Email is required"],
unique: true,
lowercase: true,
trim: true,
match: [/^\S+@\S+\.\S+$/, "Email is invalid"],
},
name: {
first: { type: String, required: true, trim: true },
last: { type: String, trim: true },
},
passwordHash: { type: String, required: true, select: false },
role: {
type: String,
enum: ["reader", "author", "admin"],
default: "reader",
},
bio: { type: String, maxlength: 500 },
loginCount: { type: Number, default: 0, min: 0 },
lastLoginAt: Date,
},
{ timestamps: true },
);
A few things worth noticing:
- Nested objects like
namedefine subpaths (name.first,name.last) without creating a separate document. select: falsehidespasswordHashfrom query results unless you explicitly ask for it with.select("+passwordHash"). That's a simple, effective guard against leaking hashes in API responses.timestamps: trueaddscreatedAtandupdatedAtand keepsupdatedAtcurrent on saves and updates.unique: trueis not a validator. It tells Mongoose to build a unique index. Duplicate violations come back from the server as anE11000error, not as a MongooseValidationError.
Schema Types
The core types are String, Number, Boolean, Date, Buffer, ObjectId, Array, Map, Decimal128, Mixed, BigInt, and UUID. Arrays are written as [Type] or [subSchema]:
const postSchema = new Schema(
{
title: { type: String, required: true, trim: true, maxlength: 200 },
slug: { type: String, required: true, unique: true },
body: { type: String, required: true },
author: { type: Schema.Types.ObjectId, ref: "User", required: true },
tags: { type: [String], default: [] },
status: {
type: String,
enum: ["draft", "published", "archived"],
default: "draft",
},
publishedAt: Date,
comments: [
new Schema(
{
author: { type: Schema.Types.ObjectId, ref: "User" },
text: { type: String, required: true, maxlength: 2000 },
},
{ timestamps: true },
),
],
stats: {
views: { type: Number, default: 0 },
likes: { type: Number, default: 0 },
},
},
{ timestamps: true },
);
The comments array uses a subdocument schema, so each comment gets its own _id, validation, and timestamps. Arrays of subdocuments are great for small, bounded lists. For comments that can grow into the thousands, a separate collection is usually the better design, since documents have a 16 MB limit and large arrays make every read of the post heavier.
Indexes in the Schema
You can declare compound indexes on the schema itself:
postSchema.index({ status: 1, publishedAt: -1 });
postSchema.index({ author: 1, createdAt: -1 });
postSchema.index({ tags: 1 });
Mongoose calls createIndexes() for each model on startup when autoIndex is enabled (the default). That's convenient in development, but in production you may want to disable it and manage indexes through migrations, since building an index on a large collection is not something you want triggered by a deploy.
mongoose.set("autoIndex", process.env.NODE_ENV !== "production");
Validation
Mongoose runs validation before save() and create(). Built-in validators include required, min, max, minlength, maxlength, enum, and match. For anything else, write a custom validator:
postSchema.path("tags").validate({
validator: (tags) => tags.length <= 10,
message: "A post can have at most 10 tags",
});
postSchema.path("publishedAt").validate({
validator: function (value) {
// `this` is the document during save validation
return this.status !== "published" || value != null;
},
message: "Published posts need a publishedAt date",
});
Validation failures throw a ValidationError with details per path:
try {
await Post.create({ title: "", body: "Hi", author: userId, slug: "hi" });
} catch (err) {
if (err instanceof mongoose.Error.ValidationError) {
console.log(Object.keys(err.errors)); // [ 'title' ]
console.log(err.errors.title.message); // Path `title` is required.
}
}
Validation on Updates
Here's the first big surprise. Update queries do not run validators by default. This passes, even though role has an enum:
await User.updateOne({ email: "ada@example.com" }, { role: "superuser" });
Turn on update validators per query, or globally:
await User.updateOne(
{ email: "ada@example.com" },
{ role: "superuser" },
{ runValidators: true },
);
// ValidationError: `superuser` is not a valid enum value for path `role`.
mongoose.set("runValidators", true); // apply to all update queries
Even with runValidators, update validators only check the paths in the update, and this inside a custom validator refers to the query, not a document. Validators that compare two fields, like the publishedAt example, need special handling in updates.
Models: Compiling a Schema
A model is a class compiled from a schema and bound to a collection:
export const User = mongoose.models.User ?? mongoose.model("User", userSchema);
export const Post = mongoose.models.Post ?? mongoose.model("Post", postSchema);
Mongoose pluralizes and lowercases the model name to get the collection (User becomes users). The mongoose.models.User ?? guard prevents an OverwriteModelError when a module is evaluated more than once, which happens with hot reloading in frameworks like Next.js.
Models give you static query methods (find, findById, updateOne, deleteMany, aggregate) and create document instances:
const ada = await User.create({
email: "Ada@Example.com ",
name: { first: "Ada", last: "Lovelace" },
passwordHash: await hash("correct horse battery staple"),
role: "author",
});
console.log(ada.email); // ada@example.com
const post = new Post({
title: "Notes on the Analytical Engine",
slug: "notes-analytical-engine",
body: "...",
author: ada._id,
});
await post.save();
const recent = await Post.find({ status: "published" })
.sort({ publishedAt: -1 })
.limit(10)
.populate("author", "name")
.lean();
Use .lean() for read-only results you're sending straight to a client. You skip document hydration and get plain objects, which is noticeably faster on larger result sets.
Virtuals, Methods, and Statics
Schemas can carry behavior, not just structure.
Virtuals
A virtual is a computed property that isn't stored:
userSchema.virtual("fullName").get(function () {
return [this.name.first, this.name.last].filter(Boolean).join(" ");
});
userSchema.set("toJSON", { virtuals: true });
Virtuals aren't included in JSON.stringify output unless you enable them in toJSON, and they aren't available on .lean() results at all, since lean skips the document layer.
Instance Methods
Instance methods are available on each document:
import bcrypt from "bcrypt";
userSchema.methods.verifyPassword = function (plain) {
return bcrypt.compare(plain, this.passwordHash);
};
const user = await User.findOne({ email }).select("+passwordHash");
const ok = user && (await user.verifyPassword(password));
Statics
Statics are attached to the model itself, which makes them a clean home for reusable queries:
postSchema.statics.findPublishedByTag = function (tag, limit = 20) {
return this.find({ status: "published", tags: tag })
.sort({ publishedAt: -1 })
.limit(limit);
};
const jsPosts = await Post.findPublishedByTag("javascript").lean();
Use regular function syntax, not arrow functions, for methods, statics, virtuals, and middleware. Mongoose binds this to the document or model, and arrow functions ignore that binding.
Middleware
Middleware (also called hooks) lets you run code before or after specific operations. Mongoose has four kinds, and knowing which kind you're writing is the key to getting them right:
| Type | Runs on | this refers to |
|---|---|---|
| Document | save, validate, deleteOne (document), updateOne (doc) | The document |
| Query | find, findOne, updateOne, updateMany, deleteMany, findOneAndUpdate, etc. | The query |
| Aggregate | aggregate | The aggregation |
| Model | insertMany, bulkWrite | The model |
Document Middleware: Hashing a Password
The classic use case:
userSchema.pre("save", async function () {
if (!this.isModified("passwordHash")) return;
this.passwordHash = await bcrypt.hash(this.passwordHash, 12);
});
In this pattern the controller sets passwordHash to the plain password and the hook hashes it on save. The isModified check is essential: without it, every save of the user (updating their bio, for instance) would hash the already-hashed value again, and they'd never log in again.
Since Mongoose 5, you can write async middleware and skip the next callback entirely. Throwing an error inside the hook aborts the operation.
Document Middleware: Deriving Fields
import slugify from "slugify";
postSchema.pre("validate", function () {
if (this.isModified("title") && !this.slug) {
this.slug = slugify(this.title, { lower: true, strict: true });
}
if (
this.isModified("status") &&
this.status === "published" &&
!this.publishedAt
) {
this.publishedAt = new Date();
}
});
Using pre('validate') rather than pre('save') means the generated slug exists before the required validator checks it.
Query Middleware: Soft Deletes
Query middleware is ideal for rules that should apply to every read. Here's a soft-delete filter:
postSchema.add({ deletedAt: { type: Date, default: null } });
postSchema.pre(/^find/, function () {
if (this.getOptions().withDeleted) return;
this.where({ deletedAt: null });
});
// Normal queries exclude deleted posts
await Post.find({ author: ada._id });
// Opt in when you need them
await Post.find({ author: ada._id }).setOptions({ withDeleted: true });
The regex /^find/ matches find, findOne, findOneAndUpdate, and the other find* operations. Inside query middleware, this is the Query object, so you modify it with methods like where(), select(), and getFilter(). Note that countDocuments and aggregate aren't covered by this hook, which is a common source of mismatched counts. The broader pattern is covered in Implementing Soft Deletes in MongoDB.
Post Middleware: Reacting After the Fact
post hooks run after the operation and receive the result:
postSchema.post("save", function (doc) {
if (doc.status === "published") {
searchIndexQueue.add({ postId: doc._id.toString() });
}
});
post hooks are also a clean place to translate low-level errors. This one converts duplicate key errors into a friendlier message:
function handleDuplicate(error, res, next) {
if (error.name === "MongoServerError" && error.code === 11000) {
next(
new Error(
`Duplicate value for ${Object.keys(error.keyValue).join(", ")}`,
),
);
} else {
next(error);
}
}
userSchema.post("save", handleDuplicate);
userSchema.post("findOneAndUpdate", handleDuplicate);
Error-handling middleware is identified by having three parameters (error, res, next).
Aggregate Middleware
If you add a soft-delete query hook, add a matching aggregate hook so reports stay consistent:
postSchema.pre("aggregate", function () {
this.pipeline().unshift({ $match: { deletedAt: null } });
});
Registering Middleware Before Compiling
Hooks must be added to the schema before you call mongoose.model(). Middleware added afterward is silently ignored. Keeping schema definition, hooks, and model compilation in the same file, in that order, avoids this entirely.
The Middleware Coverage Gap
This is the most important thing to internalize: document middleware only runs when you work with documents. Here's what that means in practice:
// Runs pre('save') hooks: password gets hashed
const user = await User.findById(id);
user.passwordHash = "new-password";
await user.save();
// Does NOT run pre('save'): password stored in plain text!
await User.updateOne({ _id: id }, { passwordHash: "new-password" });
// Does NOT run pre('save') either
await User.findByIdAndUpdate(id, { passwordHash: "new-password" });
If logic must run on every change to a field, either force all changes through save(), or add query middleware for the update operations too:
userSchema.pre(["updateOne", "findOneAndUpdate"], async function () {
const update = this.getUpdate();
const target = update.$set ?? update;
if (target.passwordHash) {
target.passwordHash = await bcrypt.hash(target.passwordHash, 12);
}
});
That works, but it's clearly more fragile than the document version. In practice, the cleanest approach is a service function like changePassword(userId, plain) that loads the document and calls save(), and a code review rule that password updates go through it.
Common Pitfalls
Using arrow functions in hooks and methods. Arrow functions don't get Mongoose's this binding. Use function () {}.
Expecting update queries to validate. They don't unless you pass runValidators: true or set it globally.
Relying on save hooks for updates. updateOne, updateMany, and findOneAndUpdate skip document middleware. Decide which operations your logic must cover and hook those explicitly.
Treating unique as validation. It creates an index. Handle E11000 errors separately, and make sure the index actually exists in production if autoIndex is off.
Adding hooks after mongoose.model(). They won't run. Define everything on the schema first.
Unbounded subdocument arrays. Embedding comments or logs that grow forever leads to huge documents and slow reads. Move growing lists into their own collection.
Forgetting that lean results skip everything document-related. No virtuals, no getters, no instance methods. That's the point of .lean(), but it catches people when a fullName field suddenly goes missing from API responses.
Conclusion
Mongoose's three layers each have a clear job. Schemas define types, defaults, validation, and indexes. Models turn schemas into a query API with statics for reusable queries and instance methods for per-document behavior. Middleware runs your logic at defined points, with document hooks operating on this document and query hooks operating on this query. Most Mongoose surprises come from mixing those up, particularly expecting save hooks and validators to run on update queries.
Pick one model in your project and audit it: list every place that modifies it, and check whether each code path actually triggers the validators and hooks you're relying on. If some don't, either add the matching query middleware or funnel those changes through save().


