Type something to search...
Understanding BSON: How MongoDB Stores Your Data

Understanding BSON: How MongoDB Stores Your Data

When you type { name: "Ada", age: 36 } into mongosh, it's natural to assume MongoDB stores that JSON somewhere. It doesn't. The shell and every driver convert your document into BSON (Binary JSON) before sending it over the wire, and BSON is what lives in memory and on disk. Most of the time you never notice. Then one day a query for { price: 10 } finds nothing even though you can see a price of 10 in the document, and you discover that one was a string.

BSON exists because JSON is a great interchange format but a mediocre database format. It has no date type, only one kind of number, no binary data, and to find a field you have to parse the whole text. BSON adds a richer type system and a layout that's fast for machines to traverse, at the cost of not being human-readable.

This guide covers how BSON is structured byte by byte, the types it supports and when to use each, how numbers and dates behave across drivers, Extended JSON, how type affects comparison and sorting, and the practical storage implications of the format.

Why Not Just Store JSON?

JSON has six value types: string, number, boolean, null, object, and array. That's not enough for a database:

  • Dates. JSON has no date type, so dates become strings or numbers, and the database can't know which fields are dates.
  • Numbers. JSON has a single "number" type with no distinction between integers, floating point, or exact decimals. A database needs to know whether 0.1 + 0.2 should equal 0.3.
  • Binary data. Images, hashes, encrypted values, and UUIDs have to be base64-encoded into strings.
  • Speed. To read one field from a JSON document, you parse from the start until you find it. There are no length markers to skip over parts you don't care about.

BSON fixes all four. It adds types like Date, ObjectId, 32-bit and 64-bit integers, Decimal128, and binary data, and it prefixes documents and strings with their lengths so a parser can jump over whole sections without reading them.

The Byte-Level Structure

A BSON document has a simple shape:

document  ::= int32 (total size in bytes)  element*  0x00
element   ::= type byte  field name (C string)  value

Every document starts with a 4-byte little-endian integer giving its total size, then a list of elements, then a terminating null byte. Each element is a 1-byte type code, the field name as a null-terminated string, and the value encoded according to its type.

Here's { "hello": "world" } in BSON, all 22 bytes of it:

16 00 00 00                   total size: 22 bytes
02                            type 0x02: string
68 65 6c 6c 6f 00             field name "hello" + null
06 00 00 00                   string length: 6 (including null)
77 6f 72 6c 64 00             "world" + null
00                            end of document

That length prefix is what makes BSON traversable. If a document has a large embedded object in the first field and you want the tenth field, the parser reads the embedded document's size and skips it entirely. The same applies to strings.

You can see this for yourself in Node.js with the bson package (the same library the MongoDB driver uses):

import { BSON } from "bson";

const bytes = BSON.serialize({ hello: "world" });
console.log(bytes.length); // 22
console.log(Buffer.from(bytes).toString("hex"));
// 160000000268656c6c6f0006000000776f726c640000

console.log(BSON.deserialize(bytes)); // { hello: 'world' }

Or in Python, where the bson module ships with PyMongo:

import bson

data = bson.encode({"hello": "world"})
print(len(data))       # 22
print(data.hex())      # 160000000268656c6c6f0006000000776f726c640000
print(bson.decode(data))

Arrays Are Documents Too

BSON doesn't have a separate array layout. An array is encoded as an embedded document whose field names are the indexes as strings: "0", "1", "2", and so on. So { tags: ["a", "b"] } is stored roughly like { tags: { "0": "a", "1": "b" } } with the type byte marked as array (0x04) instead of document (0x03). This is invisible in practice, but it explains why array indexes appear as dotted paths like tags.0 in queries.

The BSON Types You'll Actually Use

The spec defines around twenty types, several of them deprecated. These are the ones that matter day to day:

TypeCodeAlias (for $type)SizeNotes
Double0x01"double"8 bytes64-bit IEEE 754 floating point
String0x02"string"4 + len + 1UTF-8, length-prefixed
Object0x03"object"variableEmbedded document
Array0x04"array"variableStored as a document with numeric-string keys
Binary0x05"binData"4 + 1 + lenHas a subtype byte (generic, UUID, encrypted, and others)
ObjectId0x07"objectId"12 bytesDefault _id type, contains a timestamp
Boolean0x08"bool"1 byte
Date0x09"date"8 bytesMilliseconds since the Unix epoch, UTC
Null0x0A"null"0 bytes
Regex0x0B"regex"variablePattern and options
Int320x10"int"4 bytes
Timestamp0x11"timestamp"8 bytesInternal replication type, not for application dates
Int640x12"long"8 bytes
Decimal1280x13"decimal"16 bytesExact base-10 decimal, ideal for money
MinKey/MaxKey0xFF/0x7F"minKey"/"maxKey"0 bytesCompare lower/higher than every other value

A few deserve extra attention.

ObjectId

An ObjectId is 12 bytes: a 4-byte timestamp (seconds since the epoch), a 5-byte random value unique to the process, and a 3-byte incrementing counter. That design means ids can be generated on the client without coordinating with the server, and they sort roughly by creation time. You can pull the timestamp back out:

ObjectId("66f8a1c2d3e4f5a6b7c8d9e0").getTimestamp();
// ISODate('2024-09-29T00:39:30.000Z')

Date vs. Timestamp

Date is the type you want for application data: a signed 64-bit count of milliseconds since January 1, 1970 UTC. It carries no time zone; it's always UTC, and conversion to local time happens in your application or in aggregation operators that accept a timezone argument.

Timestamp looks similar but is an internal type used by replication (the oplog). It's seconds plus an ordinal counter. Don't use it for "createdAt" fields.

The most common date mistake isn't choosing the wrong type, though. It's storing dates as strings. "2026-09-27" sorts correctly only if every value uses the exact same format, and you can't use date operators like $dateTrunc on it without converting first.

Binary and UUIDs

Binary values carry a subtype byte. Subtype 0 is generic bytes, and subtype 4 is the standard UUID representation. Subtype 3 is a legacy UUID format that different drivers encoded with different byte orders, which caused years of cross-language headaches. If you store UUIDs, make sure every driver in your stack uses the standard subtype 4 (in PyMongo, set uuidRepresentation="standard" on the client; the Node.js driver uses the standard form with its UUID class). Newer subtypes exist for encrypted fields and for compact vector storage used by vector search.

Numbers: The Biggest Source of Surprises

BSON has four numeric types, and the one you get depends on which tool wrote the value.

In mongosh, a plain number literal is a double, whether or not it has a decimal point:

db.demo.insertOne({
  a: 5,
  b: 5.5,
  c: NumberInt(5),
  d: NumberLong(5),
  e: NumberDecimal("5.50"),
});

db.demo.aggregate([
  {
    $project: {
      _id: 0,
      a: { $type: "$a" },
      b: { $type: "$b" },
      c: { $type: "$c" },
      d: { $type: "$d" },
      e: { $type: "$e" },
    },
  },
]);
[ { a: 'double', b: 'double', c: 'int', d: 'long', e: 'decimal' } ]

In the Node.js driver, integers that fit in 32 bits are serialized as int32 by default, other numbers as double. For explicit control, use the wrapper classes from the driver: Int32, Long, Double, and Decimal128.

In PyMongo, a Python int becomes int32 or int64 depending on its size, a float becomes a double, and you use bson.decimal128.Decimal128 for decimals.

The good news: MongoDB compares numeric types by value, so a query for { qty: 5 } matches an int32 5, an int64 5, and a double 5.0. Mixed numeric types rarely break queries. They do matter for:

  • Precision. Doubles can't represent most decimal fractions exactly. 0.1 + 0.2 is 0.30000000000000004. For money, use Decimal128, covered in depth in Storing Money and Decimals in MongoDB with Decimal128.
  • Large integers. JavaScript numbers lose precision above 2^53. If you store 64-bit ids (from Twitter-style systems or other databases), read them as Long or strings, not plain numbers.
  • Schema validation. A validator requiring bsonType: "int" rejects a double, even 5.0. Use bsonType: "number" to accept any numeric type.

Extended JSON: BSON in Text Form

Because BSON has types JSON lacks, you need a way to write BSON as text without losing information, for exports, logs, and APIs. That's Extended JSON (EJSON). It represents special types as objects with $-prefixed keys.

There are two modes. Canonical mode preserves every type exactly. Relaxed mode is more readable and uses native JSON numbers where it's safe:

{
  "_id": { "$oid": "66f8a1c2d3e4f5a6b7c8d9e0" },
  "createdAt": { "$date": "2026-09-27T06:21:00Z" },
  "views": { "$numberLong": "9007199254740993" },
  "price": { "$numberDecimal": "19.99" },
  "rating": 4.5
}

mongoexport writes relaxed Extended JSON by default, and mongoimport reads it back with full type fidelity. In code, the bson package exposes EJSON:

import { EJSON, ObjectId } from "bson";

const doc = { _id: new ObjectId(), at: new Date("2026-09-27T06:21:00Z") };

console.log(EJSON.stringify(doc, { relaxed: true }));
// {"_id":{"$oid":"66f..."},"at":{"$date":"2026-09-27T06:21:00Z"}}

const back = EJSON.parse('{"at":{"$date":"2026-09-27T06:21:00Z"}}');
console.log(back.at instanceof Date); // true

This matters when you build APIs. If you send a document through JSON.stringify, an ObjectId becomes a plain string and a Decimal128 becomes an object like { "$numberDecimal": "19.99" }, and neither round-trips cleanly. Decide deliberately how each type should appear in your API responses.

How Types Affect Queries and Sorting

Queries match by type bracket. A filter on a string value doesn't match a number, even if they look the same:

db.products.insertMany([
  { sku: "A1", price: 10 },
  { sku: "A2", price: "10" },
]);

db.products.find({ price: 10 }); // only A1
db.products.find({ price: "10" }); // only A2

To find bad data, query by type:

db.products.find({ price: { $type: "string" } });

And to fix it, use an update with an aggregation pipeline and $toDouble or $toDecimal:

db.products.updateMany({ price: { $type: "string" } }, [
  { $set: { price: { $toDecimal: "$price" } } },
]);

When sorting a field that holds mixed types, MongoDB uses a fixed cross-type order. From lowest to highest:

  1. MinKey
  2. Null
  3. Numbers (int, long, double, decimal, compared by value)
  4. Symbol, String
  5. Object
  6. Array
  7. BinData
  8. ObjectId
  9. Boolean
  10. Date
  11. Timestamp
  12. Regular Expression
  13. MaxKey

So if some documents have price: null and others have numbers, the nulls sort first in ascending order. Missing fields are treated like null for sorting.

Field Order Is Part of Equality

Because a BSON document is an ordered list of elements, field order matters when you compare whole embedded documents:

db.places.insertOne({ loc: { city: "Paris", country: "FR" } });

db.places.find({ loc: { city: "Paris", country: "FR" } }); // matches
db.places.find({ loc: { country: "FR", city: "Paris" } }); // no match!
db.places.find({ "loc.city": "Paris", "loc.country": "FR" }); // matches, any order

Exact embedded-document matches are brittle. Use dot notation on individual fields instead.

Storage Implications

BSON's design has a few practical effects on how much space your data uses.

Field names are stored in every document. A collection of ten million documents, each with a field called customerLifetimeValueEstimate, stores that 29-byte name ten million times. WiredTiger's block compression (snappy by default) recovers much of this on disk, but names still take room in memory and in network traffic. Clear names are worth it; absurdly long ones aren't.

BSON can be larger than JSON. Length prefixes, type bytes, and fixed-width numbers add overhead. A small integer like 7 is one character in JSON and four or eight bytes in BSON. In exchange, you get fast traversal and exact types.

Documents max out at 16 MB. That limit is on the BSON size, not the JSON text. You can check a document's size with the $bsonSize aggregation operator:

db.orders.aggregate([
  { $project: { size: { $bsonSize: "$$ROOT" } } },
  { $sort: { size: -1 } },
  { $limit: 5 },
]);

This is a great way to find documents with runaway arrays before they become a problem.

Common Mistakes

Storing numbers and dates as strings. It's the most common BSON-related bug. Form input and CSV imports arrive as strings, and it's easy to write them straight into the database. Convert at the boundary, and consider schema validation to reject the wrong types.

Assuming mongosh numbers are integers. A literal 5 in the shell is a double. If a field must be an integer (for validation, or for consistency with an application that writes int32), use NumberInt() or NumberLong().

Using double for money. Rounding errors accumulate. Use Decimal128, or store integer minor units (cents) as int64.

Mixing UUID representations. Two services writing UUIDs with different binary subtypes produce values that never match each other. Standardize on subtype 4.

Matching embedded documents exactly. Field order and extra fields both break exact matches. Query individual fields with dot notation.

Serializing BSON types with plain JSON.stringify. Types get flattened or mangled. Use EJSON or explicit conversions in your API layer.

Conclusion

BSON is the binary format behind every MongoDB document. Length prefixes make it fast to traverse, and its type system adds what JSON lacks: dates, multiple numeric types, exact decimals, ObjectIds, and binary data. Those types aren't just storage details. They decide which queries match, how values sort, what schema validation accepts, and how data survives a trip through Extended JSON.

For a next step, run db.yourCollection.aggregate([{ $group: { _id: { $type: "$someField" }, count: { $sum: 1 } } }]) against a field in one of your own collections. If you see more than one type come back, you've found your first data cleanup task.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading