Type something to search...
Storing Files in MongoDB with GridFS

Storing Files in MongoDB with GridFS

Sooner or later, most applications need to store files: profile photos, PDF invoices, CSV exports, audio clips. If MongoDB is already your primary database, it's tempting to put the files there too. The first attempt usually looks like stuffing the bytes into a Binary field on a document, and that works right up until someone uploads a 20 MB video and the insert fails because a single BSON document can't exceed 16 MB.

GridFS is MongoDB's answer to that limit. It's a specification, implemented by every official driver, that splits a file into small chunks, stores each chunk as its own document, and keeps a separate metadata document describing the whole file. Because the file is spread across many documents, there's no practical size limit, and because each chunk is independent, you can stream a file in and out without loading it all into memory or even read just a byte range from the middle.

This guide covers how GridFS stores data, how to upload and download files with the Node.js and Python drivers, how to serve files over HTTP with range support, how to manage metadata, and when you should skip GridFS and use object storage instead.

How GridFS Stores a File

GridFS uses two collections per bucket. By default the bucket is named fs, so you get:

  • fs.files: one document per file, holding the filename, total length, chunk size, upload date, and any custom metadata.
  • fs.chunks: many documents per file, each holding a slice of the binary data.

A file document looks like this:

{
  _id: ObjectId("66f8a1c2e4b0a1b2c3d4e5f6"),
  length: 5242880,
  chunkSize: 261120,
  uploadDate: ISODate("2026-09-23T06:03:00Z"),
  filename: "q3-report.pdf",
  metadata: { ownerId: "u_1042", contentType: "application/pdf" }
}

And each chunk references it through files_id, with n giving its position:

{
  _id: ObjectId("66f8a1c2e4b0a1b2c3d4e5f7"),
  files_id: ObjectId("66f8a1c2e4b0a1b2c3d4e5f6"),
  n: 0,
  data: BinData(0, "JVBERi0xLjcKJ...")
}

The default chunk size is 255 KiB (261,120 bytes). It's chosen to be comfortably below the 16 MB document limit while keeping the number of chunks reasonable. A 5 MB file becomes 21 chunk documents: 20 full chunks and a final partial one.

When a driver reads a file, it looks up the file document, then queries fs.chunks for all documents with that files_id, sorted by n, and concatenates the data fields. That query is fast because drivers create two indexes the first time you write to a bucket:

db.fs.files.getIndexes();
// [ { key: { _id: 1 } }, { key: { filename: 1, uploadDate: 1 } } ]

db.fs.chunks.getIndexes();
// [ { key: { _id: 1 } }, { key: { files_id: 1, n: 1 }, unique: true } ]

The unique index on { files_id: 1, n: 1 } guarantees there's only ever one chunk at each position, and it makes range reads cheap: to start reading at byte 10,000,000, the driver computes the chunk number and seeks directly to it.

Uploading Files with Node.js

In the Node.js driver (6.x), GridFS lives in the GridFSBucket class. You create a bucket from a Db, then pipe a readable stream into an upload stream.

import { createReadStream } from "node:fs";
import { pipeline } from "node:stream/promises";
import { MongoClient, GridFSBucket } from "mongodb";

const client = new MongoClient(process.env.MONGODB_URI);
await client.connect();

const db = client.db("media");
const bucket = new GridFSBucket(db, { bucketName: "uploads" });

const uploadStream = bucket.openUploadStream("q3-report.pdf", {
  metadata: { ownerId: "u_1042", contentType: "application/pdf" },
});

await pipeline(createReadStream("./q3-report.pdf"), uploadStream);

console.log("Stored file with id", uploadStream.id.toString());

A few things to notice:

  • bucketName: "uploads" gives you uploads.files and uploads.chunks instead of the default fs.* collections. Use separate buckets when you have different kinds of files with different retention or access rules (for example, avatars and invoices).
  • The upload stream writes chunks as data arrives, so memory usage stays around one chunk regardless of the file size.
  • The file document in uploads.files is inserted after the last chunk is written. If the upload fails halfway, you're left with orphaned chunks and no file document, which is why a cleanup job is worth having (more on that later).

Older driver versions accepted contentType and aliases as top-level options. Those fields are deprecated by the GridFS spec and removed in current drivers, so put the MIME type inside metadata as shown above.

If you want to control the file's _id (for example, to use a UUID or an ID you've already stored on another document), use openUploadStreamWithId:

import { ObjectId } from "mongodb";

const fileId = new ObjectId();
const stream = bucket.openUploadStreamWithId(fileId, "avatar.webp", {
  chunkSizeBytes: 64 * 1024,
  metadata: { userId: "u_1042", contentType: "image/webp" },
});

The chunkSizeBytes option overrides the chunk size for this file only. Smaller chunks help when you'll often read small ranges; larger chunks mean fewer documents for big files. For most workloads, the default is fine.

Downloading and Streaming Files

Reading is the mirror image: open a download stream and pipe it somewhere.

import { createWriteStream } from "node:fs";

await pipeline(
  bucket.openDownloadStream(fileId),
  createWriteStream("./downloaded.pdf"),
);

You can also look files up by name. Because GridFS allows several files with the same filename, openDownloadStreamByName takes a revision option: -1 (the default) means the most recent upload, 0 means the oldest.

const latest = bucket.openDownloadStreamByName("q3-report.pdf");
const original = bucket.openDownloadStreamByName("q3-report.pdf", {
  revision: 0,
});

This makes a simple versioning scheme almost free. Upload the new version under the same name, and readers get the newest one automatically.

Serving Files Over HTTP

The most common use of GridFS is serving uploads from a web server. Here's an Express route that sets the right headers and supports HTTP range requests, which browsers use for video seeking and resumable downloads:

import express from "express";
import { ObjectId } from "mongodb";

const app = express();

app.get("/files/:id", async (req, res, next) => {
  try {
    if (!ObjectId.isValid(req.params.id)) return res.sendStatus(400);
    const _id = new ObjectId(req.params.id);

    const [file] = await bucket.find({ _id }).limit(1).toArray();
    if (!file) return res.sendStatus(404);

    const type = file.metadata?.contentType ?? "application/octet-stream";
    res.set("Content-Type", type);
    res.set("Accept-Ranges", "bytes");
    res.set("ETag", `"${file._id}"`);

    const range = req.headers.range?.match(/^bytes=(\d*)-(\d*)$/);
    if (range) {
      const start = range[1] ? Number(range[1]) : 0;
      const end = range[2] ? Number(range[2]) : file.length - 1;
      if (start >= file.length || end < start) {
        res.set("Content-Range", `bytes */${file.length}`);
        return res.sendStatus(416);
      }
      res.status(206);
      res.set("Content-Range", `bytes ${start}-${end}/${file.length}`);
      res.set("Content-Length", String(end - start + 1));
      // GridFS "end" is exclusive, HTTP ranges are inclusive
      return bucket
        .openDownloadStream(_id, { start, end: end + 1 })
        .on("error", next)
        .pipe(res);
    }

    res.set("Content-Length", String(file.length));
    bucket.openDownloadStream(_id).on("error", next).pipe(res);
  } catch (err) {
    next(err);
  }
});

The start and end options tell the driver to skip straight to the chunk containing start and stop after end. Note the off-by-one: GridFS treats end as exclusive, while the HTTP Range header is inclusive.

Because file documents never change after upload, using the _id as an ETag is safe, and you can add long-lived Cache-Control headers too. In production, put a CDN in front of this route so repeated requests don't hit your database at all.

Working with GridFS in Python

PyMongo ships a gridfs package with the same bucket API. The method names follow Python conventions:

from pymongo import MongoClient
from gridfs import GridFSBucket

client = MongoClient("mongodb://localhost:27017")
db = client["media"]
bucket = GridFSBucket(db, bucket_name="uploads")

# Upload from an open file object
with open("q3-report.pdf", "rb") as f:
    file_id = bucket.upload_from_stream(
        "q3-report.pdf",
        f,
        metadata={"ownerId": "u_1042", "contentType": "application/pdf"},
    )

# Download to a file object
with open("downloaded.pdf", "wb") as out:
    bucket.download_to_stream(file_id, out)

# Or read it in pieces
grid_out = bucket.open_download_stream(file_id)
print(grid_out.filename, grid_out.length)
header = grid_out.read(1024)

open_download_stream returns a GridOut object, which behaves like a read-only file: you can read(), seek(), and iterate over it. That's handy for tasks like reading the first few kilobytes of a file to sniff its type.

If you're using PyMongo's native async API (AsyncMongoClient, available since PyMongo 4.9), the gridfs package also provides an async bucket class with awaitable versions of these methods. Motor has its own GridFS implementation, but Motor is deprecated in favor of the async PyMongo API, so new code should use the latter.

Querying and Managing File Metadata

Since uploads.files is an ordinary collection, you can query it with ordinary filters. The bucket's find() method is a thin wrapper around a find on the files collection:

const recentPdfs = await bucket
  .find(
    {
      "metadata.ownerId": "u_1042",
      "metadata.contentType": "application/pdf",
      uploadDate: { $gte: new Date("2026-09-01") },
    },
    { sort: { uploadDate: -1 }, limit: 20 },
  )
  .toArray();

for (const f of recentPdfs) {
  console.log(f.filename, (f.length / 1024 / 1024).toFixed(1), "MB");
}

If you query metadata fields often, index them just as you would on any other collection:

db.uploads.files.createIndex({ "metadata.ownerId": 1, uploadDate: -1 });

Other management operations on the bucket:

await bucket.rename(fileId, "q3-report-final.pdf"); // updates files doc only
await bucket.delete(fileId); // removes the files doc and all its chunks
await bucket.drop(); // drops both collections of the bucket

There's no built-in "update metadata" method, but you don't need one. Update the file document directly with updateOne on uploads.files. Just never modify length, chunkSize, or _id, since the driver relies on them to read the chunks correctly.

Linking Files to Your Domain Data

A good pattern is to keep business data in your own collections and store only the GridFS file ID there:

await db.collection("invoices").insertOne({
  invoiceNumber: "INV-2026-0931",
  customerId: "c_77",
  total: 1249.0,
  pdfFileId: uploadStream.id,
  createdAt: new Date(),
});

This keeps your application queries fast (they never touch uploads.chunks), and it lets you decide per collection how files relate to records. When you delete an invoice, delete its file too, ideally in the same code path.

Using mongofiles from the Command Line

The MongoDB Database Tools include mongofiles, a small CLI for GridFS. It's useful for quick inspections, scripts, and moving a few files around:

# Upload a file to the default "fs" bucket
mongofiles --uri="mongodb://localhost:27017/media" put ./logo.png

# List files (optionally filtered by a filename prefix)
mongofiles --uri="mongodb://localhost:27017/media" list

# Download a file by name
mongofiles --uri="mongodb://localhost:27017/media" get logo.png

# Use a custom bucket with --prefix
mongofiles --uri="mongodb://localhost:27017/media" --prefix=uploads list

The list command prints each filename and its size in bytes:

connected to: mongodb://localhost:27017/media
logo.png 48213
q3-report.pdf 5242880

Cleaning Up Orphaned Chunks

Because a file document is only written after all of its chunks, a crashed upload can leave chunks with no parent. Deletes can leave the opposite problem if the process dies between removing the file document and removing the chunks. Neither case corrupts anything, but orphaned chunks waste storage.

An aggregation can find chunk groups whose files_id has no matching file document:

db.uploads.chunks.aggregate([
  { $group: { _id: "$files_id", chunks: { $sum: 1 } } },
  {
    $lookup: {
      from: "uploads.files",
      localField: "_id",
      foreignField: "_id",
      as: "file",
    },
  },
  { $match: { file: { $size: 0 } } },
  { $project: { chunks: 1 } },
]);

The $group stage scans every chunk, so on large buckets run it off-peak or against a secondary. Once you have the orphaned IDs, remove them with deleteMany({ files_id: { $in: ids } }). To avoid deleting chunks of uploads that are still in progress, only remove orphans whose ObjectId timestamp is older than an hour or so.

GridFS vs. Object Storage

GridFS is a solid tool, but it isn't always the right one. Object stores like Amazon S3, Google Cloud Storage, or Azure Blob Storage are purpose-built for files and are usually cheaper per gigabyte.

ConcernGridFSObject storage (S3 and similar)
Extra infrastructureNone, uses your existing clusterA separate service and credentials
Backup and replicationIncluded with your database backupsConfigured separately
Storage costDatabase disk (often expensive)Cheap, with tiered storage classes
Serving to browsersThrough your app serverDirect via pre-signed URLs or a CDN
Consistency with your dataSame cluster, same backupsNeeds its own cleanup and reconciliation
Effect on the working setChunks compete for WiredTiger cacheNo impact on database memory

Pick GridFS when:

  • You want files replicated, backed up, and restored together with the data that references them.
  • You're running on-premises or in an environment where adding an object store is a hassle.
  • Files are moderate in size and volume, and you need range reads without a separate service.
  • You want to query file metadata with the full power of MongoDB queries.

Pick object storage when:

  • You're storing a large volume of media (terabytes of images or video).
  • Files should be served directly to clients through a CDN.
  • Your database is on a managed tier where storage is priced much higher than blob storage.

A common hybrid approach is to store the file in S3 and keep a metadata document in MongoDB with the object key, size, checksum, and owner. You get cheap storage plus rich querying.

Best Practices

Don't use GridFS for small files by default. If your files are reliably under a few megabytes (thumbnails, small JSON blobs), storing them as a Binary field in a regular document is simpler and faster: one document read instead of two queries. Reach for GridFS when files can exceed the 16 MB limit or when you need streaming and range reads.

Treat files as immutable. GridFS has no "update file contents" operation, and trying to rewrite chunks in place is asking for trouble. Upload a new file, point your domain document at the new ID, then delete the old one.

Keep file traffic off the primary when you can. Large downloads read a lot of data. If slightly stale reads are acceptable, a secondaryPreferred read preference on the bucket's database spreads that load, which is covered in more depth in MongoDB Read Preferences and Write Concerns Explained.

Validate before you store. Check file size and MIME type on upload, and never trust the client-supplied filename for anything security-sensitive. Store the original name in metadata and use the _id in URLs.

Remember that chunks count toward your working set. Heavy file reads can push hot application data out of the WiredTiger cache. If file serving is a big part of your workload, consider a dedicated database or cluster for GridFS buckets, or put a CDN in front.

Shard on files_id if you shard at all. If a bucket grows large enough to need sharding, shard the chunks collection on { files_id: 1, n: 1 } so each file's chunks stay together. The files collection is usually small enough to leave unsharded.

Conclusion

GridFS removes the 16 MB ceiling by splitting files into 255 KiB chunks across two collections, and every official driver gives you a streaming API for uploads, downloads, and byte-range reads. Add metadata inside the metadata field, index what you query, link file IDs from your domain documents, and schedule a small cleanup job for orphaned chunks. For very large media libraries, object storage plus a metadata document in MongoDB is usually the better fit.

As a next step, take one upload endpoint in your app, switch it to openUploadStream with a pipeline(), and add the range-aware download route above. Then test it by seeking through a video in the browser and watching the 206 Partial Content responses in your network tab.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading