Type something to search...
Building a GraphQL API with MongoDB and Apollo Server

Building a GraphQL API with MongoDB and Apollo Server

GraphQL and MongoDB look like they were designed for each other. A GraphQL query describes a nested tree of fields, and a MongoDB document is a nested tree of fields. Clients ask for exactly the shape they need, and your resolvers pull it out of documents that already look a lot like the response.

The fit is real, but the details matter. A naive resolver setup turns one GraphQL query for twenty books and their authors into twenty-one database queries. IDs need translating between GraphQL strings and ObjectId. Pagination needs a design that doesn't collapse at page 500. And because clients can compose arbitrary queries, you need limits so one request can't ask for the entire database.

This guide builds a book catalog API with Apollo Server and the official MongoDB Node.js driver. It covers the schema, resolvers, a per-request context, batching with DataLoader, cursor-based pagination, mutations with validation, error handling, and the protections a public GraphQL endpoint needs.

Project Setup

mkdir graphql-books && cd graphql-books
npm init -y
npm install @apollo/server graphql mongodb dataloader

Set "type": "module" in package.json so you can use ES modules and top-level await. The examples use the @apollo/server package (Apollo Server 4 and later) and its built-in standalone HTTP server.

The data model has two collections:

// authors
{ _id: ObjectId("..."), name: "Ursula K. Le Guin", born: 1929 }

// books
{
  _id: ObjectId("..."),
  title: "The Left Hand of Darkness",
  authorId: ObjectId("..."),
  year: 1969,
  genres: ["sci-fi"],
  rating: 4.4
}

Books reference authors by authorId, a classic one-to-many relationship. The GraphQL API will expose it as a nested author field on Book and a books field on Author.

Defining the Schema

The GraphQL schema is your API contract. Write it in SDL (Schema Definition Language):

// src/schema.js
export const typeDefs = `#graphql
  type Author {
    id: ID!
    name: String!
    born: Int
    books(limit: Int = 10): [Book!]!
  }

  type Book {
    id: ID!
    title: String!
    year: Int
    genres: [String!]!
    rating: Float
    author: Author
  }

  type BookEdge {
    cursor: String!
    node: Book!
  }

  type PageInfo {
    endCursor: String
    hasNextPage: Boolean!
  }

  type BookConnection {
    edges: [BookEdge!]!
    pageInfo: PageInfo!
  }

  input BookFilter {
    genre: String
    authorId: ID
    minYear: Int
  }

  input AddBookInput {
    title: String!
    authorId: ID!
    year: Int
    genres: [String!]
  }

  type Query {
    book(id: ID!): Book
    books(filter: BookFilter, first: Int = 20, after: String): BookConnection!
    author(id: ID!): Author
  }

  type Mutation {
    addBook(input: AddBookInput!): Book!
    rateBook(id: ID!, rating: Float!): Book
    deleteBook(id: ID!): Boolean!
  }
`;

A few design choices here:

  • IDs are strings (ID!), not ObjectId. GraphQL clients don't know about BSON. Resolvers convert at the boundary.
  • The books list uses a connection type with edges and cursors. It's more verbose than a plain array, but it supports stable pagination and is the convention most GraphQL clients (like Relay and Apollo Client) understand.
  • Mutations take input types. AddBookInput groups arguments and gives you a single place to extend later.

Connecting to MongoDB

As with any Node.js service, create one client for the process:

// src/db.js
import { MongoClient } from "mongodb";

export const client = new MongoClient(process.env.MONGODB_URI, {
  appName: "graphql-books",
  serverSelectionTimeoutMS: 5000,
});

export const db = client.db(process.env.MONGODB_DB ?? "library");
export const books = db.collection("books");
export const authors = db.collection("authors");

export async function connectDb() {
  await client.connect();
  await books.createIndex({ authorId: 1, year: -1 });
  await books.createIndex({ genres: 1, _id: 1 });
}

The { authorId: 1 } index is not optional. Every Author.books resolver and every DataLoader batch filters on it.

The N+1 Problem

Before writing resolvers, it's worth seeing the trap. Consider this query:

query {
  books(first: 20) {
    edges {
      node {
        title
        author {
          name
        }
      }
    }
  }
}

A straightforward Book.author resolver would call authors.findOne({ _id: book.authorId }) once per book. That's one query for the list plus twenty queries for authors: the N+1 problem. On a page of 100 books, it's 101 round trips.

DataLoader fixes this by collecting every load(id) call made during one tick of the event loop, then calling your batch function once with all the IDs. Twenty author lookups become a single find({ _id: { $in: [...] } }).

Per-Request Loaders

DataLoader caches results, so loaders must be created per request. A loader shared across requests would serve stale data and could leak data between users.

// src/loaders.js
import DataLoader from "dataloader";
import { authors, books } from "./db.js";

export function createLoaders() {
  return {
    authorById: new DataLoader(
      async (ids) => {
        const docs = await authors.find({ _id: { $in: ids } }).toArray();
        const byId = new Map(docs.map((d) => [d._id.toHexString(), d]));
        return ids.map((id) => byId.get(id.toHexString()) ?? null);
      },
      { cacheKeyFn: (id) => id.toHexString() },
    ),

    booksByAuthorId: new DataLoader(
      async (authorIds) => {
        const docs = await books
          .find({ authorId: { $in: authorIds } })
          .sort({ year: -1 })
          .toArray();
        const grouped = new Map();
        for (const doc of docs) {
          const key = doc.authorId.toHexString();
          if (!grouped.has(key)) grouped.set(key, []);
          grouped.get(key).push(doc);
        }
        return authorIds.map((id) => grouped.get(id.toHexString()) ?? []);
      },
      { cacheKeyFn: (id) => id.toHexString() },
    ),
  };
}

Two rules make DataLoader work correctly:

  1. The batch function must return results in the same order as the input keys, with one entry per key. MongoDB's $in returns documents in no particular order, which is why the code builds a Map and then maps over the original IDs.
  2. Keys need a stable identity. Two ObjectId instances with the same value are different objects, so cacheKeyFn converts them to hex strings for caching and deduplication.

Resolvers

Resolvers map schema fields to data. Start with a helper that converts GraphQL IDs into ObjectId values and rejects garbage with a proper GraphQL error:

// src/resolvers.js
import { ObjectId } from "mongodb";
import { GraphQLError } from "graphql";
import { books } from "./db.js";

function toObjectId(id, field = "id") {
  if (typeof id !== "string" || !/^[0-9a-f]{24}$/i.test(id)) {
    throw new GraphQLError(`Invalid ${field}`, {
      extensions: { code: "BAD_USER_INPUT", field },
    });
  }
  return new ObjectId(id);
}

const encodeCursor = (id) =>
  Buffer.from(id.toHexString()).toString("base64url");
const decodeCursor = (cursor) =>
  toObjectId(Buffer.from(cursor, "base64url").toString(), "after");

Query Resolvers

export const resolvers = {
  Query: {
    book: (_, { id }) => books.findOne({ _id: toObjectId(id) }),

    author: (_, { id }, { loaders }) =>
      loaders.authorById.load(toObjectId(id)),

    books: async (_, { filter = {}, first, after }) => {
      const limit = Math.min(Math.max(first, 1), 100);
      const query = {};

      if (filter.genre) query.genres = filter.genre;
      if (filter.authorId) query.authorId = toObjectId(filter.authorId, "authorId");
      if (filter.minYear != null) query.year = { $gte: filter.minYear };
      if (after) query._id = { $gt: decodeCursor(after) };

      const docs = await books
        .find(query)
        .sort({ _id: 1 })
        .limit(limit + 1)
        .toArray();

      const hasNextPage = docs.length > limit;
      const page = hasNextPage ? docs.slice(0, limit) : docs;

      return {
        edges: page.map((doc) => ({ cursor: encodeCursor(doc._id), node: doc })),
        pageInfo: {
          endCursor: page.length ? encodeCursor(page.at(-1)._id) : null,
          hasNextPage,
        },
      };
    },
  },

The books resolver uses range-based pagination on _id. It fetches one extra document to know whether there's a next page, without running a separate count. Unlike skip, this stays fast no matter how deep a client paginates, because each page is an index seek. The first-argument clamp keeps a client from requesting 10,000 books at once.

Field Resolvers

Field resolvers handle the id mapping and the relationships:

  Book: {
    id: (book) => book._id.toHexString(),
    genres: (book) => book.genres ?? [],
    author: (book, _, { loaders }) =>
      book.authorId ? loaders.authorById.load(book.authorId) : null,
  },

  Author: {
    id: (author) => author._id.toHexString(),
    books: async (author, { limit }, { loaders }) => {
      const all = await loaders.booksByAuthorId.load(author._id);
      return all.slice(0, Math.min(limit, 50));
    },
  },

The first argument to a field resolver is the parent object, which here is the raw MongoDB document returned by the parent resolver. That's why Book.id reads book._id. Fields that match by name (title, year, rating) don't need resolvers at all; Apollo's default resolver reads the property directly.

Mutations

  Mutation: {
    addBook: async (_, { input }, { user, loaders }) => {
      requireUser(user);

      const title = input.title.trim();
      if (!title || title.length > 300) {
        throw new GraphQLError("Title must be 1-300 characters", {
          extensions: { code: "BAD_USER_INPUT", field: "title" },
        });
      }

      const authorId = toObjectId(input.authorId, "authorId");
      if (!(await loaders.authorById.load(authorId))) {
        throw new GraphQLError("Author not found", {
          extensions: { code: "BAD_USER_INPUT", field: "authorId" },
        });
      }

      const doc = {
        title,
        authorId,
        year: input.year ?? null,
        genres: input.genres ?? [],
        rating: null,
        createdAt: new Date(),
      };
      const { insertedId } = await books.insertOne(doc);
      return { _id: insertedId, ...doc };
    },

    rateBook: async (_, { id, rating }, { user }) => {
      requireUser(user);
      if (rating < 0 || rating > 5) {
        throw new GraphQLError("Rating must be between 0 and 5", {
          extensions: { code: "BAD_USER_INPUT", field: "rating" },
        });
      }
      return books.findOneAndUpdate(
        { _id: toObjectId(id) },
        { $set: { rating } },
        { returnDocument: "after" },
      );
    },

    deleteBook: async (_, { id }, { user }) => {
      requireUser(user, "admin");
      const { deletedCount } = await books.deleteOne({ _id: toObjectId(id) });
      return deletedCount === 1;
    },
  },
};

With driver 6.x, findOneAndUpdate returns the document itself (or null), which maps neatly to the nullable Book return type of rateBook.

GraphQL validates argument types for you: rating is guaranteed to be a number and input.title a string. What it can't check is business rules like ranges and lengths, so those live in the resolver. For larger schemas, a validation library or custom scalars keep this tidy.

The Auth Helper

function requireUser(user, role) {
  if (!user) {
    throw new GraphQLError("Not authenticated", {
      extensions: { code: "UNAUTHENTICATED", http: { status: 401 } },
    });
  }
  if (role && user.role !== role) {
    throw new GraphQLError("Not allowed", {
      extensions: { code: "FORBIDDEN", http: { status: 403 } },
    });
  }
}

Starting the Server

The context function runs once per request. It's where you authenticate the caller and create fresh loaders:

// src/index.js
import { ApolloServer } from "@apollo/server";
import { startStandaloneServer } from "@apollo/server/standalone";
import { typeDefs } from "./schema.js";
import { resolvers } from "./resolvers.js";
import { connectDb } from "./db.js";
import { createLoaders } from "./loaders.js";
import { verifyToken } from "./auth.js";

await connectDb();

const server = new ApolloServer({
  typeDefs,
  resolvers,
  introspection: process.env.NODE_ENV !== "production",
});

const { url } = await startStandaloneServer(server, {
  listen: { port: Number(process.env.PORT ?? 4000) },
  context: async ({ req }) => ({
    user: await verifyToken(req.headers.authorization),
    loaders: createLoaders(),
  }),
});

console.log(`GraphQL ready at ${url}`);

verifyToken stands in for your auth logic (for example, verifying a JWT and returning { id, role } or null). If you already have an Express app, Apollo provides integration packages that mount the server as middleware instead; pick the one matching your Apollo Server and Express versions.

Trying It Out

Open the URL in a browser to use Apollo's sandbox, or send a query with curl:

curl -s localhost:4000/ \
  -H "Content-Type: application/json" \
  -d '{"query":"{ books(first: 2, filter: { genre: \"sci-fi\" }) { edges { cursor node { title author { name } } } pageInfo { hasNextPage endCursor } } }"}'
{
  "data": {
    "books": {
      "edges": [
        {
          "cursor": "NjZmMGExYjJjM2Q0ZTVmNmE3YjhjOWQw",
          "node": {
            "title": "The Left Hand of Darkness",
            "author": { "name": "Ursula K. Le Guin" }
          }
        },
        {
          "cursor": "NjZmMGExYjJjM2Q0ZTVmNmE3YjhjOWQx",
          "node": {
            "title": "The Dispossessed",
            "author": { "name": "Ursula K. Le Guin" }
          }
        }
      ],
      "pageInfo": {
        "hasNextPage": true,
        "endCursor": "NjZmMGExYjJjM2Q0ZTVmNmE3YjhjOWQx"
      }
    }
  }
}

Both books share an author, and thanks to DataLoader's deduplication, that author was fetched once, in one query. If you enable MongoDB's profiler or command monitoring on the client, you'll see exactly two queries for this request: one for books and one for authors.

Fetching Only What's Requested

GraphQL clients choose their fields, but the resolvers above still fetch whole documents. For small documents that's fine. When documents carry large fields (long descriptions, embedded reviews, content blobs), fetching them for a query that only wants titles wastes bandwidth and memory.

You can read the requested fields from the resolver's fourth argument, info, and build a projection. Libraries such as graphql-parse-resolve-info make this easier. Keep it simple, though: always include _id and any fields that child resolvers depend on (like authorId), or nested fields will silently resolve to null.

In many APIs, a simpler approach works just as well: exclude a known list of heavy fields from list queries with a fixed projection, and fetch them only in the single-item resolver.

Protecting a Public Endpoint

A REST endpoint does one fixed thing. A GraphQL endpoint does whatever the client composes, which makes a few protections essential:

  • Limit query depth. A query like author { books { author { books { ... } } } } can nest arbitrarily. Use a validation rule such as graphql-depth-limit to reject queries deeper than, say, 7 levels.
  • Cap list sizes in every resolver, as the first and limit clamps above do. Never trust a client-supplied page size.
  • Consider cost analysis for public APIs, rejecting queries whose estimated cost (fields times list sizes) exceeds a budget.
  • Disable introspection in production if your API isn't meant to be public, and consider persisted queries so only known operations run.
  • Set a server-side timeout on expensive operations with maxTimeMS so a slow query can't pile up connections.
  • Mask internal errors. Apollo reports unexpected errors as INTERNAL_SERVER_ERROR. Use the formatError hook to strip stack traces and driver messages in production.

Common Pitfalls

Resolving relationships with findOne per parent. This is the N+1 problem. Batch with DataLoader or fetch relationships with $lookup in the parent resolver.

Sharing DataLoaders across requests. Loaders cache per instance. Create them in the context function so each request starts fresh.

Returning results in the wrong order from a batch function. DataLoader maps results to keys by position. Always reorder to match the input keys.

Leaking ObjectId into the API. Convert to strings in id field resolvers and parse incoming IDs with validation. An unvalidated new ObjectId(input) throws on bad input and turns a client mistake into a server error.

Offset pagination on large collections. skip gets slower the deeper you go. Cursor pagination on an indexed field stays constant.

Missing indexes on foreign keys. Every authorId lookup needs an index, or each batched query becomes a collection scan.

Conclusion

A MongoDB-backed GraphQL API comes together cleanly once you respect a few rules. Keep the schema in terms of string IDs and connection types, convert to ObjectId at the resolver boundary, and resolve relationships through per-request DataLoaders so nested queries cost a handful of round trips instead of hundreds. Add cursor-based pagination on indexed fields, business-rule validation in mutations, and depth and size limits for anything public.

To go further, add a reviews collection with a Book.reviews field resolved through a new DataLoader, and watch the query count with command monitoring as you nest it. Seeing the batching work in your own logs is the best way to make the pattern stick. If you need a refresher on modeling the relationship itself, One-to-Many Relationships in MongoDB covers the trade-offs.

Tags :
Share :

Related Posts

A Complete Guide to MongoDB Query Operators

A Complete Guide to MongoDB Query Operators

Your first MongoDB queries are usually simple equality filters: find the user with this email, find orders with this status. That covers a surprising

Continue Reading
Async MongoDB in Python with Motor and FastAPI

Async MongoDB in Python with Motor and FastAPI

FastAPI runs your endpoints on an event loop. That's what lets a single worker juggle hundreds of concurrent requests: while one request waits on the

Continue Reading
Atlas Online Archive: Tiering Cold Data to Cut Costs

Atlas Online Archive: Tiering Cold Data to Cut Costs

Look at almost any production database and you'll find the same shape. A small slice of recent data gets nearly all the reads and writes: this week's

Continue Reading