Type something to search...
Connecting Python Apps to MongoDB with PyMongo

Connecting Python Apps to MongoDB with PyMongo

Python dictionaries and MongoDB documents are close cousins. A document is a set of named fields that can hold strings, numbers, lists, and nested objects, which is exactly what a dict is. That makes MongoDB one of the most natural databases to use from Python: you insert dictionaries, you query with dictionaries, and you get dictionaries back.

PyMongo is MongoDB's official Python driver, and it's what almost every Python MongoDB tool builds on. It handles connection pooling, replica set discovery, automatic retries, and conversion between Python types and BSON. Recent versions also include a native async API, which replaces the older Motor library for asyncio applications.

This guide covers installing PyMongo, connecting safely, CRUD operations, cursors and projections, working with ObjectId and dates, indexes, bulk writes, transactions, error handling, type hints, and the async client. The examples use PyMongo 4.x with Python 3.11 or later.

Installing PyMongo

Create a virtual environment and install the driver:

python -m venv .venv
source .venv/bin/activate
pip install pymongo

Recent PyMongo versions include dnspython as a dependency, which is required for mongodb+srv:// connection strings (the format Atlas uses). If you're on an old version that complains about SRV lookups, upgrade rather than installing extras.

Verify the installation:

python -c "import pymongo; print(pymongo.version)"

Note that you should not install the separate bson package from PyPI. PyMongo ships its own bson module, and the unrelated package with the same name breaks it in confusing ways.

Connecting

The central object is MongoClient:

# connect.py
import os
from pymongo import MongoClient

client = MongoClient(
    os.environ["MONGODB_URI"],
    appname="inventory-service",
    serverSelectionTimeoutMS=5000,
)

print(client.admin.command("ping"))
$ MONGODB_URI="mongodb://localhost:27017" python connect.py
{'ok': 1.0}

Creating a MongoClient returns immediately and starts connecting in background threads. The first operation (here, ping) waits until a suitable server is available, or raises an error once serverSelectionTimeoutMS has elapsed. The default is 30 seconds, which is long enough that an unreachable database makes your app look frozen, so a shorter value is usually better.

Connection Strings

Keep the URI in an environment variable, never in code:

# Local
export MONGODB_URI="mongodb://localhost:27017"

# Atlas (SRV format)
export MONGODB_URI="mongodb+srv://inventory_app:s3cret@cluster0.abcd1.mongodb.net/?retryWrites=true&w=majority&appName=inventory-service"

If the password contains special characters like @ or :, percent-encode it with urllib.parse.quote_plus() before building the URI.

One Client per Process

MongoClient is thread-safe and maintains a connection pool. Create it once and share it. A common pattern is a small module:

# db.py
import os
from pymongo import MongoClient

client = MongoClient(
    os.environ["MONGODB_URI"],
    appname="inventory-service",
    serverSelectionTimeoutMS=5000,
    tz_aware=True,
)
db = client[os.environ.get("MONGODB_DB", "inventory")]

Every module that needs the database imports db from here. Python caches modules after the first import, so everyone shares the same client.

One important caveat: MongoClient is not fork-safe. If you use a pre-forking server like Gunicorn with preload_app, or multiprocessing, create the client in each worker after the fork, not in the parent. PyMongo warns you if it detects a client being used across a fork.

Databases and Collections

Databases and collections are accessed with attribute or dictionary syntax, and both are created lazily on first write:

from db import db

products = db.products        # attribute style
orders = db["orders"]         # dictionary style, needed for names like "order-items"

CRUD Operations

Inserting

from datetime import datetime, timezone

result = products.insert_one({
    "sku": "LAMP-001",
    "name": "Brass Desk Lamp",
    "price": 39.99,
    "stock": 12,
    "tags": ["lighting", "office"],
    "created_at": datetime.now(timezone.utc),
})
print(result.inserted_id)
# 66f5c8e2a1b2c3d4e5f60718

products.insert_many([
    {"sku": "CHAIR-01", "name": "Oak Chair", "price": 189.0, "stock": 4, "tags": ["furniture"]},
    {"sku": "MUG-07", "name": "Stoneware Mug", "price": 14.5, "stock": 0, "tags": ["kitchen"]},
])

PyMongo generates an _id for you and, as a side effect, adds it to the dictionary you passed in. If you insert the same dict object twice in a loop, the second insert fails with a duplicate key error because it now carries the first _id. Build a fresh dict each time.

Reading

find_one() returns a dict or None. find() returns a Cursor you can iterate:

lamp = products.find_one({"sku": "LAMP-001"})
print(lamp["name"])  # Brass Desk Lamp

in_stock = products.find(
    {"stock": {"$gt": 0}, "price": {"$lt": 200}},
    projection={"_id": 0, "sku": 1, "name": 1, "price": 1},
).sort("price", -1).limit(10)

for p in in_stock:
    print(p)
{'sku': 'CHAIR-01', 'name': 'Oak Chair', 'price': 189.0}
{'sku': 'LAMP-001', 'name': 'Brass Desk Lamp', 'price': 39.99}

Query operators like $gt are written as dictionary keys, exactly as in mongosh. For multi-field sorts, pass a list of tuples: .sort([("stock", -1), ("name", 1)]).

A cursor fetches results lazily in batches and can only be iterated once. If you need a list, call list(cursor), but only for bounded queries. For large result sets, iterate directly so you're not holding everything in memory.

Updating

from pymongo import ReturnDocument

result = products.update_one(
    {"sku": "LAMP-001", "stock": {"$gt": 0}},
    {"$inc": {"stock": -1}, "$set": {"updated_at": datetime.now(timezone.utc)}},
)
print(result.matched_count, result.modified_count)  # 1 1

products.update_many({"tags": "kitchen"}, {"$addToSet": {"tags": "gift-ideas"}})

updated = products.find_one_and_update(
    {"sku": "MUG-07"},
    {"$set": {"stock": 30}},
    return_document=ReturnDocument.AFTER,
)
print(updated["stock"])  # 30

The update_one call above is an atomic, safe decrement: it only matches when stock is positive, so two concurrent orders can't take the stock below zero. That pattern (condition in the filter, operator in the update) is much safer than reading a document, changing it in Python, and writing it back. The full set of update operators is covered in MongoDB Update Operators: $set, $inc, $push, $pull, and More.

For upserts, pass upsert=True, and use replace_one when you want to swap the whole document.

Deleting and Counting

products.delete_one({"sku": "MUG-07"})
products.delete_many({"stock": 0, "tags": "discontinued"})

print(products.count_documents({"stock": {"$gt": 0}}))  # exact, filtered
print(products.estimated_document_count())               # fast, metadata-based

Cursor.count() was removed in PyMongo 4, so older code that calls it needs updating to count_documents.

ObjectId and Dates

ObjectId

_id values are ObjectId instances from the bson module. Querying with the string form returns nothing:

from bson import ObjectId
from bson.errors import InvalidId

def get_product(product_id: str):
    try:
        oid = ObjectId(product_id)
    except InvalidId:
        return None
    return products.find_one({"_id": oid})

When returning documents from a web API, convert ObjectId to str, since the standard json module can't serialize it. bson.json_util.dumps() is an alternative that produces MongoDB Extended JSON.

Datetimes

BSON stores dates as UTC milliseconds. By default, PyMongo returns naive datetime objects representing UTC, which is an easy way to introduce timezone bugs. Passing tz_aware=True to MongoClient (as in the db.py module above) makes it return timezone-aware UTC datetimes instead:

doc = products.find_one({"sku": "LAMP-001"})
print(doc["created_at"])  # 2026-09-26 07:59:12.345000+00:00

Always store timezone-aware UTC datetimes, and convert to local time only for display. Note that BSON dates have millisecond precision, so microseconds are truncated on the way in.

Aggregation

aggregate() takes a list of stages and returns a cursor:

pipeline = [
    {"$unwind": "$tags"},
    {"$group": {"_id": "$tags", "products": {"$sum": 1}, "avg_price": {"$avg": "$price"}}},
    {"$sort": {"products": -1}},
    {"$limit": 5},
]

for row in products.aggregate(pipeline):
    print(row)
{'_id': 'lighting', 'products': 3, 'avg_price': 54.66}
{'_id': 'furniture', 'products': 2, 'avg_price': 214.5}

Pipelines are plain Python lists of dicts, so you can build them conditionally: append a $match stage only when a filter is set, for example.

Indexes

from pymongo import ASCENDING, DESCENDING

products.create_index("sku", unique=True)
products.create_index([("tags", ASCENDING), ("price", DESCENDING)])
orders.create_index("created_at", expireAfterSeconds=60 * 60 * 24 * 90)

print(products.index_information())

create_index is idempotent when the definition matches an existing index, so running it at startup is safe for small collections. For large collections, create indexes as a deliberate deployment step.

Bulk Writes

When you're importing or syncing many documents, send them in batches with bulk_write:

from pymongo import UpdateOne

feed = [
    {"sku": "LAMP-001", "price": 42.0, "stock": 20},
    {"sku": "DESK-02", "price": 349.0, "stock": 3},
]

ops = [
    UpdateOne(
        {"sku": item["sku"]},
        {"$set": {"price": item["price"], "stock": item["stock"]}},
        upsert=True,
    )
    for item in feed
]

result = products.bulk_write(ops, ordered=False)
print(result.matched_count, result.upserted_count)

ordered=False lets the server continue past individual failures and process operations in parallel, which is usually what you want for imports.

Transactions

Multi-document transactions require a replica set or sharded cluster (Atlas clusters qualify). The with_transaction helper handles commit, retries on transient errors, and aborts on exceptions:

def place_order(session, sku: str, qty: int, customer_id):
    result = products.update_one(
        {"sku": sku, "stock": {"$gte": qty}},
        {"$inc": {"stock": -qty}},
        session=session,
    )
    if result.modified_count == 0:
        raise ValueError(f"Insufficient stock for {sku}")

    orders.insert_one(
        {"sku": sku, "qty": qty, "customer_id": customer_id,
         "created_at": datetime.now(timezone.utc)},
        session=session,
    )

with client.start_session() as session:
    session.with_transaction(lambda s: place_order(s, "LAMP-001", 2, customer_id))

Every operation inside the transaction must receive the session. Forgetting it on one call silently runs that operation outside the transaction, which is the most common transaction bug in any driver.

Error Handling

PyMongo's exceptions live in pymongo.errors. These are the ones you'll handle most often:

from pymongo.errors import (
    DuplicateKeyError,
    ServerSelectionTimeoutError,
    OperationFailure,
)

try:
    products.insert_one({"sku": "LAMP-001", "name": "Duplicate Lamp"})
except DuplicateKeyError as e:
    print("SKU already exists:", e.details.get("keyValue"))
except ServerSelectionTimeoutError:
    print("Database unreachable")
except OperationFailure as e:
    print("Server error", e.code, e.details)

DuplicateKeyError is a subclass of OperationFailure, so order the except clauses from most to least specific. Retryable writes are enabled by default, so a single transient network error during a failover is retried automatically before you ever see an exception.

Type Hints

PyMongo supports generics, so you can describe your documents with TypedDict and get editor help:

import os
from typing import TypedDict, NotRequired
from bson import ObjectId
from pymongo import MongoClient
from pymongo.collection import Collection

class Product(TypedDict):
    _id: NotRequired[ObjectId]
    sku: str
    name: str
    price: float
    stock: int
    tags: list[str]

client: MongoClient[dict] = MongoClient(os.environ["MONGODB_URI"])
typed_products: Collection[Product] = client.inventory.products

item = typed_products.find_one({"sku": "LAMP-001"})
if item:
    print(item["price"] * 2)  # the type checker knows price is a float

As with any type hints, this is static checking only. MongoDB can still return documents that don't match, so validate data at the edges if it matters. Server-side schema validation is covered in Schema Validation in MongoDB with JSON Schema.

The Async API

For asyncio applications (FastAPI, aiohttp, async workers), PyMongo now includes a native async client, introduced in PyMongo 4.9 and matured over subsequent releases:

import asyncio
import os
from pymongo import AsyncMongoClient

async def main():
    client = AsyncMongoClient(os.environ["MONGODB_URI"], tz_aware=True)
    products = client.inventory.products

    await products.insert_one({"sku": "LAMP-002", "name": "Floor Lamp", "price": 89.0, "stock": 5})

    async for p in products.find({"price": {"$lt": 100}}).sort("price", 1):
        print(p["name"])

    total = await products.count_documents({})
    print("total:", total)

    await client.close()

asyncio.run(main())

The API mirrors the synchronous one, with await on operations and async for on cursors. If you have an existing app using Motor, the older async driver, be aware that Motor is deprecated in favor of this API and reaches end of life in 2026. Migrating is mostly a matter of changing the import and client class, then fixing a few method differences (such as close() being awaitable). Async MongoDB in Python with Motor and FastAPI walks through a full async app.

Don't mix the two: the synchronous MongoClient blocks the event loop when used inside async code.

Common Mistakes

Creating a client per request. MongoClient owns a connection pool and background monitoring threads. Create it once per process and reuse it.

Querying _id with a string. Convert with ObjectId(value) and catch InvalidId for bad input.

Reusing a dict across inserts. insert_one adds _id to your dict, so a second insert of the same object fails as a duplicate.

Naive datetimes. Store aware UTC datetimes and use tz_aware=True so you get aware ones back.

Sharing a client across fork(). Create clients in each worker process after forking.

Using the sync client in async code. Use AsyncMongoClient in asyncio apps, and plan a migration off Motor.

Installing the bson package from PyPI. It conflicts with PyMongo's built-in bson module. Uninstall it if it's there.

Conclusion

PyMongo makes MongoDB feel native in Python: documents are dicts, queries are dicts, and results come back as dicts. Using it well comes down to a few habits: one shared MongoClient with a sensible server selection timeout, atomic update operators instead of read-modify-write, explicit ObjectId conversion, timezone-aware datetimes, and sessions passed to every operation inside a transaction. For asyncio code, the built-in AsyncMongoClient is now the way forward.

Your next step: create a db.py module like the one in this guide, point it at a local or Atlas cluster, and port one existing script to use it with tz_aware=True. Then run count_documents and a small aggregation against real data to get a feel for how directly Python maps onto MongoDB.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading