
Building a URL Shortener with Node.js and MongoDB
A URL shortener looks like a weekend toy: take a long URL, hand back a short code, redirect when someone visits it. But once you think about it for more than a minute, it turns into a surprisingly good tour of real database design. Codes have to be unique without a central counter becoming a bottleneck. Redirects have to be fast because they sit in front of every click. Links may need to expire. And the moment anyone uses it, they want to know how many clicks they got and from where.
MongoDB fits this problem well. A short code makes a natural _id, unique indexes handle collisions for free, TTL indexes expire links without a cron job, $inc counts clicks atomically, and a time-series collection stores click events efficiently for analytics.
This guide builds a working shortener with Express 5 and the official Node.js driver: the data model, generating codes, handling custom aliases, the redirect path, expiration, click analytics with aggregation, and what to change when it needs to scale.
Project Setup
Start a new project and install the dependencies:
mkdir shorty && cd shorty
npm init -y
npm pkg set type=module
npm install express mongodb zod
You'll need a MongoDB server. A local instance, a Docker container (docker run -d -p 27017:27017 mongo:8.0), or a free Atlas M0 cluster all work. Put the connection string in an environment variable:
export MONGODB_URI="mongodb://localhost:27017"
export BASE_URL="http://localhost:3000"
The Data Model
Two collections cover everything:
links: one document per short link.clicks: one document per redirect, used for analytics.
A link document looks like this:
{
_id: "k3Xp9Qa", // the short code
url: "https://example.com/some/very/long/path?utm_source=newsletter",
ownerId: "user_123",
createdAt: ISODate("2026-09-24T08:52:00Z"),
expiresAt: ISODate("2026-10-24T08:52:00Z"), // optional
clicks: 0
}
Using the code as _id is the key decision. The _id index already exists and is unique, so lookups by code are a single index hit and collisions are rejected by the database automatically. There's no need for a separate code field and a second unique index.
Click events are small and append-only, which is exactly what time-series collections are built for:
{
ts: ISODate("2026-09-24T09:14:03Z"),
meta: { code: "k3Xp9Qa" },
referrer: "news.ycombinator.com",
country: "PT",
userAgent: "Mozilla/5.0 ..."
}
Connecting and Creating Indexes
Create one client for the whole app and set up the collections on startup:
// db.js
import { MongoClient } from "mongodb";
export const client = new MongoClient(process.env.MONGODB_URI);
export const db = client.db("shorty");
export const links = db.collection("links");
export const clicks = db.collection("clicks");
export async function initDb() {
await client.connect();
// Expire links automatically once expiresAt passes
await links.createIndex({ expiresAt: 1 }, { expireAfterSeconds: 0 });
await links.createIndex({ ownerId: 1, createdAt: -1 });
const existing = await db.listCollections({ name: "clicks" }).toArray();
if (existing.length === 0) {
await db.createCollection("clicks", {
timeseries: {
timeField: "ts",
metaField: "meta",
granularity: "seconds",
},
expireAfterSeconds: 60 * 60 * 24 * 365, // keep one year of click data
});
}
}
The TTL index with expireAfterSeconds: 0 tells MongoDB to delete each link when the clock passes its expiresAt value. Documents without an expiresAt field are never deleted, so permanent links just omit it. The TTL monitor runs roughly once a minute, so expired links can linger briefly; we'll handle that in the redirect code. See TTL indexes for more on how the background deletion works.
Generating Short Codes
You have three broad options for codes:
| Approach | Pros | Cons |
|---|---|---|
| Random base62 string | No coordination, unguessable, easy | Must handle rare collisions |
| Counter + base62 encoding | Shortest codes, no collisions | Sequential and guessable; counter is a hot spot |
| Hash of the URL | Same URL always gets same code | Collisions between different URLs; no per-user links |
Random codes are the best default. With 62 possible characters and 7 characters, there are about 3.5 trillion combinations. Even with 10 million links stored, the chance a new code collides is roughly 1 in 350,000, and when it does, the unique _id index tells you and you try again.
Use crypto.randomInt so codes are unpredictable (Math.random is not suitable for this):
// codes.js
import { randomInt } from "node:crypto";
const ALPHABET =
"0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz";
export function randomCode(length = 7) {
let code = "";
for (let i = 0; i < length; i++) code += ALPHABET[randomInt(ALPHABET.length)];
return code;
}
Creating Links
The create endpoint validates input, then tries to insert with a fresh random code, retrying on the rare collision:
// routes/links.js
import { Router } from "express";
import { z } from "zod";
import { links } from "../db.js";
import { randomCode } from "../codes.js";
const router = Router();
const CreateLink = z.object({
url: z
.string()
.url()
.max(2048)
.refine(
(u) => ["http:", "https:"].includes(new URL(u).protocol),
"Only http(s) URLs",
),
alias: z
.string()
.regex(/^[A-Za-z0-9_-]{3,32}$/)
.optional(),
expiresInDays: z.number().int().min(1).max(365).optional(),
});
const RESERVED = new Set(["api", "admin", "login", "stats", "health"]);
router.post("/api/links", async (req, res) => {
const parsed = CreateLink.safeParse(req.body);
if (!parsed.success)
return res.status(400).json({ error: parsed.error.issues });
const { url, alias, expiresInDays } = parsed.data;
if (alias && RESERVED.has(alias.toLowerCase())) {
return res.status(400).json({ error: "That alias is reserved" });
}
const now = new Date();
const doc = {
url,
ownerId: req.user?.id ?? null,
createdAt: now,
clicks: 0,
};
if (expiresInDays)
doc.expiresAt = new Date(now.getTime() + expiresInDays * 86_400_000);
const maxAttempts = alias ? 1 : 5;
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
const _id = alias ?? randomCode();
try {
await links.insertOne({ _id, ...doc });
return res
.status(201)
.json({ code: _id, shortUrl: `${process.env.BASE_URL}/${_id}` });
} catch (err) {
if (err.code !== 11000) throw err;
if (alias) return res.status(409).json({ error: "Alias already taken" });
}
}
res.status(503).json({ error: "Could not allocate a code, try again" });
});
export default router;
A few details matter here:
- Only
http:andhttps:URLs. Without that check, someone can shortenjavascript:ordata:URLs and use your domain to deliver something nasty. - Custom aliases share the
_idnamespace. A taken alias produces the same duplicate key error as a random collision, and the database decides who wins if two people claim it at the same moment. No race-prone "check then insert". - Reserved words. Aliases like
apiorstatswould shadow your own routes. - Bounded retries. Five collisions in a row with 7-character codes would mean the keyspace is nearly full, so fail loudly rather than looping forever.
Try it out:
curl -s -X POST localhost:3000/api/links \
-H "content-type: application/json" \
-d '{"url":"https://www.mongodb.com/docs/manual/core/timeseries-collections/","expiresInDays":30}'
{ "code": "k3Xp9Qa", "shortUrl": "http://localhost:3000/k3Xp9Qa" }
The Redirect Path
This is the hot path: every click goes through it, so it needs to be as fast as possible. One indexed read, one atomic increment, and a fire-and-forget analytics insert:
// routes/redirect.js
import { Router } from "express";
import { links, clicks } from "../db.js";
const router = Router();
router.get("/:code", async (req, res, next) => {
const { code } = req.params;
if (!/^[A-Za-z0-9_-]{3,32}$/.test(code)) return next();
const now = new Date();
const link = await links.findOneAndUpdate(
{
_id: code,
$or: [{ expiresAt: { $exists: false } }, { expiresAt: { $gt: now } }],
},
{ $inc: { clicks: 1 } },
{ projection: { url: 1 } },
);
if (!link) return res.status(404).send("Link not found or expired");
res.redirect(302, link.url);
clicks
.insertOne({
ts: now,
meta: { code },
referrer: refererHost(req.get("referer")),
country: req.get("cf-ipcountry") ?? null,
userAgent: req.get("user-agent")?.slice(0, 300) ?? null,
})
.catch((err) => console.error("click log failed", err));
});
function refererHost(value) {
try {
return value ? new URL(value).hostname : null;
} catch {
return null;
}
}
export default router;
findOneAndUpdate does the lookup and the counter increment in a single atomic operation, so concurrent clicks never lose counts. The expiresAt condition in the filter covers the gap before the TTL monitor deletes an expired link. With driver 6.x, findOneAndUpdate returns the document itself (or null), and by default it's the document as it was before the update, which is fine since we only need the URL.
The click insert happens after the response is sent. The user doesn't wait for analytics, and if the insert fails, the redirect still worked. The country value here assumes you're behind Cloudflare, which sets the cf-ipcountry header; with a different CDN or a GeoIP lookup, swap in your own source.
301 or 302?
A 301 (permanent) redirect is cached by browsers, so repeat visits skip your server entirely. That's great for load and terrible for analytics and for changing a link's destination later. Most shorteners use 302 (or 307) so every click is counted.
Wiring Up the App
// server.js
import express from "express";
import { initDb, client } from "./db.js";
import linkRoutes from "./routes/links.js";
import redirectRoutes from "./routes/redirect.js";
import statsRoutes from "./routes/stats.js";
const app = express();
app.use(express.json({ limit: "10kb" }));
app.use(linkRoutes);
app.use(statsRoutes);
app.use(redirectRoutes);
await initDb();
const server = app.listen(3000, () => console.log("shorty listening on :3000"));
process.on("SIGTERM", async () => {
server.close();
await client.close();
});
Express 5 forwards rejected promises from async handlers to the error handler automatically, so you don't need try/catch wrappers around every route.
Click Analytics with Aggregation
With click events in a time-series collection, stats are a couple of aggregation pipelines away. Here's clicks per day for a link over the last 30 days, plus top referrers, in one request using $facet:
// routes/stats.js
import { Router } from "express";
import { links, clicks } from "../db.js";
const router = Router();
router.get("/api/links/:code/stats", async (req, res) => {
const { code } = req.params;
const link = await links.findOne(
{ _id: code },
{ projection: { url: 1, clicks: 1, createdAt: 1 } },
);
if (!link) return res.status(404).json({ error: "Not found" });
const since = new Date(Date.now() - 30 * 86_400_000);
const [stats] = await clicks
.aggregate([
{ $match: { "meta.code": code, ts: { $gte: since } } },
{
$facet: {
daily: [
{
$group: {
_id: { $dateTrunc: { date: "$ts", unit: "day" } },
clicks: { $sum: 1 },
},
},
{ $sort: { _id: 1 } },
{ $project: { _id: 0, day: "$_id", clicks: 1 } },
],
referrers: [
{
$group: {
_id: { $ifNull: ["$referrer", "direct"] },
clicks: { $sum: 1 },
},
},
{ $sort: { clicks: -1 } },
{ $limit: 10 },
],
countries: [
{ $group: { _id: "$country", clicks: { $sum: 1 } } },
{ $sort: { clicks: -1 } },
{ $limit: 10 },
],
},
},
])
.toArray();
res.json({ ...link, last30Days: stats });
});
export default router;
A typical response:
{
"_id": "k3Xp9Qa",
"url": "https://www.mongodb.com/docs/manual/core/timeseries-collections/",
"clicks": 1284,
"createdAt": "2026-09-24T08:52:00.000Z",
"last30Days": {
"daily": [
{ "day": "2026-09-24T00:00:00.000Z", "clicks": 911 },
{ "day": "2026-09-25T00:00:00.000Z", "clicks": 373 }
],
"referrers": [
{ "_id": "news.ycombinator.com", "clicks": 702 },
{ "_id": "direct", "clicks": 389 }
],
"countries": [
{ "_id": "US", "clicks": 488 },
{ "_id": "DE", "clicks": 131 }
]
}
}
The $match on meta.code and ts is efficient because time-series collections organize data into buckets by metaField and time. For a busy service, you can also add a secondary index on { "meta.code": 1, ts: -1 }.
The clicks counter on the link document is the fast total for list views. The time-series data is the detailed record for charts. Keeping both is a small, deliberate bit of denormalization.
Listing a User's Links
The { ownerId: 1, createdAt: -1 } index supports a dashboard of a user's links with range-based pagination:
router.get("/api/me/links", async (req, res) => {
const filter = { ownerId: req.user.id };
if (req.query.before)
filter.createdAt = { $lt: new Date(String(req.query.before)) };
const items = await links
.find(filter, {
projection: { url: 1, clicks: 1, createdAt: 1, expiresAt: 1 },
})
.sort({ createdAt: -1 })
.limit(25)
.toArray();
res.json({ items, next: items.at(-1)?.createdAt ?? null });
});
Abuse Prevention
Public shorteners attract spammers and phishers quickly. At a minimum:
- Rate-limit link creation per IP and per user (a middleware such as
express-rate-limitbacked by Redis or MongoDB works). - Check destinations against a threat list such as the Google Safe Browsing API before accepting them, and periodically recheck popular links.
- Block your own domain as a destination to prevent redirect loops and chains.
- Keep a
disabledflag on links so you can take one down instantly without deleting its history, and adddisabled: { $ne: true }to the redirect filter.
Scaling Up
This design goes a long way on a single replica set. When it needs to go further:
- Cache hot links. A small in-process LRU cache or Redis in front of the redirect path absorbs viral links. The trade-off is that click counts must then be batched rather than incremented per request.
- Batch click increments. Accumulate counts in memory for a few seconds and flush them with
bulkWriteof$incoperations. This turns 10,000 writes per second on one hot document into a handful. - Shard by code. If you ever outgrow one replica set, a hashed shard key on
_idspreads random codes evenly, and every redirect targets exactly one shard. See choosing a good shard key. - Read from secondaries for stats. Analytics queries can use
readPreference: "secondaryPreferred"to keep load off the primary.
Common Pitfalls
Using Math.random() for codes. It isn't cryptographically secure, so codes can be predicted, and private links can be enumerated. Use crypto.randomInt or crypto.randomBytes.
Checking if a code exists before inserting. Two requests can both see "free" and both insert. Let the unique _id index decide and handle error 11000.
Relying solely on the TTL index for expiry. TTL deletion runs periodically, not instantly. Always check expiresAt in the redirect query.
Writing analytics before responding. The user shouldn't wait for your click log. Send the redirect first and log asynchronously.
Accepting any URL scheme. javascript: and data: URLs turn your shortener into an attack vector. Allow only http and https.
Conclusion
A URL shortener is small enough to build in an afternoon and rich enough to exercise many of MongoDB's best features: a natural _id that doubles as a unique constraint, duplicate key handling instead of race-prone checks, TTL indexes for expiry, atomic $inc counters, and time-series collections with $facet for analytics.
Get the version above running locally, then add the batched click counter: keep a Map of code to count, flush it every five seconds with a single bulkWrite, and watch how much write load disappears when you hammer one link with a load-testing tool like autocannon.


