
Implementing Rate Limiting in Next.js Route Handlers
Every public endpoint will eventually get hit harder than you planned. Sometimes it's a scraper, sometimes a bot trying passwords on your login route, sometimes your own frontend stuck in a retry loop. If the endpoint sends email, calls a paid AI API, or runs an expensive database query, a few thousand extra requests can turn into a real bill or an outage.
Rate limiting caps how many requests a client can make in a given time window. Requests over the limit get a fast, cheap 429 Too Many Requests response instead of reaching your business logic. Next.js doesn't ship a built-in rate limiter, but Route Handlers make it easy to add one.
This guide covers the common rate limiting algorithms, how to identify clients, a simple in-memory limiter for single-instance apps, a Redis-backed limiter for multi-instance and serverless deployments, the response headers clients expect, a reusable wrapper, and when to move limiting into proxy.ts.
Choosing an Algorithm
All rate limiters answer the same question, "has this client made too many requests recently?", but they define "recently" differently.
| Algorithm | How it works | Pros | Cons |
|---|---|---|---|
| Fixed window | Count requests per clock window (e.g. per minute); reset at the boundary | Simplest, one counter per key | Allows bursts of up to 2x the limit around window edges |
| Sliding window | Weight the previous window's count by how much of it overlaps the current period | Smooth, cheap, no edge bursts | Slightly approximate |
| Token bucket | Each client has a bucket that refills at a steady rate; each request takes a token | Allows controlled bursts, smooth average rate | Two values to tune (capacity, refill rate) |
For most APIs, sliding window is a good default. Token bucket suits APIs where short bursts are normal (a client loading a dashboard fires ten requests at once, then goes quiet). Fixed window is fine for coarse limits like "100 signups per IP per day".
Identifying the Client
A limit applies per key. What you choose as the key matters as much as the algorithm:
- User ID for authenticated endpoints. It's the most accurate, and it follows users across networks.
- API key for public APIs with issued keys. Each key gets its own quota, and you can vary limits per plan.
- IP address for anonymous traffic like login, signup, or contact forms.
Next.js removed request.ip in version 15, so you read the IP from headers set by your host or reverse proxy:
// lib/client-ip.ts
export function getClientIp(request: Request): string {
const forwardedFor = request.headers.get("x-forwarded-for");
if (forwardedFor) {
// The left-most address is the original client.
return forwardedFor.split(",")[0].trim();
}
return request.headers.get("x-real-ip") ?? "unknown";
}
A word of caution: X-Forwarded-For is only trustworthy if a proxy you control sets or overwrites it. If your Next.js server is exposed directly to the internet, clients can send any value they like and dodge your limits. Platforms like Vercel and Cloudflare set these headers reliably; on your own servers, configure nginx or your load balancer to overwrite them.
Also keep in mind that many users can share one IP (offices, universities, mobile carriers). Keep IP-based limits generous and prefer user or API-key limits where you can.
A Simple In-Memory Limiter
If you run a single long-lived Node.js server (one container or one VM), an in-memory limiter is enough and has no dependencies. Here's a token bucket:
// lib/rate-limit/memory.ts
type Bucket = { tokens: number; updatedAt: number };
export type RateLimitResult = {
success: boolean;
limit: number;
remaining: number;
reset: number; // Unix ms when the bucket will have at least one token
};
export function createMemoryLimiter({
capacity,
refillPerSecond,
}: {
capacity: number;
refillPerSecond: number;
}) {
const buckets = new Map<string, Bucket>();
// Drop idle buckets so memory doesn't grow forever.
const fullRefillMs = (capacity / refillPerSecond) * 1000;
setInterval(() => {
const cutoff = Date.now() - fullRefillMs;
for (const [key, bucket] of buckets) {
if (bucket.updatedAt < cutoff) buckets.delete(key);
}
}, 60_000).unref();
return function limit(key: string): RateLimitResult {
const now = Date.now();
const bucket = buckets.get(key) ?? { tokens: capacity, updatedAt: now };
// Refill based on time elapsed since the last request.
const elapsedSeconds = (now - bucket.updatedAt) / 1000;
bucket.tokens = Math.min(
capacity,
bucket.tokens + elapsedSeconds * refillPerSecond,
);
bucket.updatedAt = now;
const success = bucket.tokens >= 1;
if (success) bucket.tokens -= 1;
buckets.set(key, bucket);
const msUntilNextToken = success
? 0
: ((1 - bucket.tokens) / refillPerSecond) * 1000;
return {
success,
limit: capacity,
remaining: Math.floor(bucket.tokens),
reset: now + Math.ceil(msUntilNextToken),
};
};
}
How it works:
- Each key gets a bucket that starts full (
capacitytokens). - On each request, the bucket refills according to how much time has passed, capped at
capacity. There's no background timer per key; refilling is calculated lazily. - If at least one token is available, the request is allowed and a token is removed. Otherwise it's rejected.
- A cleanup interval deletes buckets that have been idle long enough to be full again, since a full bucket is the same as no bucket.
unref()keeps the timer from holding the process open.
Use it in a Route Handler:
// app/api/search/route.ts
import { createMemoryLimiter } from "@/lib/rate-limit/memory";
import { getClientIp } from "@/lib/client-ip";
const limit = createMemoryLimiter({ capacity: 20, refillPerSecond: 0.5 });
export async function GET(request: Request) {
const result = limit(getClientIp(request));
if (!result.success) {
return Response.json(
{ error: "Too many requests" },
{
status: 429,
headers: {
"Retry-After": String(Math.ceil((result.reset - Date.now()) / 1000)),
},
},
);
}
const q = new URL(request.url).searchParams.get("q") ?? "";
return Response.json({ query: q, results: [] });
}
This allows a burst of 20 searches, then one more every two seconds.
When In-Memory Isn't Enough
The limiter's state lives in one process's memory, which breaks down when:
- You run several instances. Each one has its own counters, so the real limit becomes
limit × instances, and a load balancer spreads an attacker's requests across all of them. - You deploy to serverless. Function instances are created and destroyed constantly, and each cold start begins with empty buckets.
- You restart. Every deploy resets every counter.
For any of these, you need shared storage. Redis is the standard choice: it's fast, supports atomic increments, and expires keys automatically.
Rate Limiting with Redis
Option 1: Upstash Ratelimit
@upstash/ratelimit implements fixed window, sliding window, and token bucket on top of Redis over HTTP. Because it talks HTTP instead of holding a TCP connection, it works well in serverless functions and edge runtimes.
npm install @upstash/ratelimit @upstash/redis
// lib/rate-limit/upstash.ts
import { Ratelimit } from "@upstash/ratelimit";
import { Redis } from "@upstash/redis";
const redis = Redis.fromEnv(); // reads UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN
export const apiLimiter = new Ratelimit({
redis,
limiter: Ratelimit.slidingWindow(60, "1 m"),
prefix: "rl:api",
analytics: true,
});
export const loginLimiter = new Ratelimit({
redis,
limiter: Ratelimit.fixedWindow(5, "15 m"),
prefix: "rl:login",
});
Each limiter has its own prefix, so keys for different endpoints never collide. slidingWindow(60, "1 m") allows 60 requests in any rolling minute. The login limiter is deliberately strict: five attempts per 15 minutes per key.
Using it in a Route Handler:
// app/api/login/route.ts
import { after } from "next/server";
import { loginLimiter } from "@/lib/rate-limit/upstash";
import { getClientIp } from "@/lib/client-ip";
export async function POST(request: Request) {
const ip = getClientIp(request);
const { success, limit, remaining, reset, pending } =
await loginLimiter.limit(ip);
// Let analytics writes finish after the response is sent.
after(() => pending);
if (!success) {
return Response.json(
{ error: "Too many login attempts. Try again later." },
{
status: 429,
headers: {
"RateLimit-Limit": String(limit),
"RateLimit-Remaining": String(remaining),
"Retry-After": String(
Math.max(1, Math.ceil((reset - Date.now()) / 1000)),
),
},
},
);
}
const { email, password } = await request.json();
// ...verify credentials
return Response.json({ ok: true, email: Boolean(email && password) });
}
limit() returns everything you need: whether the request is allowed, the limit, how many requests remain, when the window resets (Unix milliseconds), and a pending promise for background work like analytics. Passing pending to after() lets Next.js finish that work after the response is sent, so it doesn't add latency, and it isn't cut off when a serverless function freezes.
For login specifically, consider limiting on two keys: the IP, and the email being attempted. The IP limit stops one machine from trying many accounts; the email limit stops a botnet from trying many passwords on one account.
Option 2: Your Own Redis
If you already run Redis, a fixed-window limiter is a few lines with the official redis client:
// lib/rate-limit/redis.ts
import { createClient } from "redis";
let client: ReturnType<typeof createClient> | undefined;
async function getRedis() {
if (!client) {
client = createClient({ url: process.env.REDIS_URL });
client.on("error", (err) => console.error("Redis error", err));
await client.connect();
}
return client;
}
export async function fixedWindowLimit(
key: string,
limit: number,
windowSeconds: number,
) {
const redis = await getRedis();
const window = Math.floor(Date.now() / 1000 / windowSeconds);
const redisKey = `rl:${key}:${window}`;
const count = await redis.incr(redisKey);
if (count === 1) {
await redis.expire(redisKey, windowSeconds);
}
const reset = (window + 1) * windowSeconds * 1000;
return {
success: count <= limit,
limit,
remaining: Math.max(0, limit - count),
reset,
};
}
The key includes the current window number, so each window gets a fresh counter. INCR is atomic, so concurrent requests on different servers can't both read "4" and both write "5". The expire call cleans up old windows automatically. If a crash happened between INCR and EXPIRE, the key would linger without a TTL; for strict guarantees, run both in a MULTI transaction or a small Lua script. The client is created lazily and reused, which matters because opening a new connection per request is slow and can exhaust Redis's connection limit.
Sending the Right Headers
A well-behaved 429 tells the client when to try again. Two headers matter most:
Retry-After: seconds until the client may retry. This is the standard HTTP header, and many HTTP clients and SDKs respect it automatically.RateLimit-Limit/RateLimit-Remaining/RateLimit-Reset: the current quota. These come from an IETF draft and are widely used, often with anX-prefix in older APIs. Sending them on successful responses too lets well-written clients slow down before they hit the limit.
Keep 429 responses small and cheap. Don't render a page, don't query the database, and don't log the full request body. The point is to spend as little as possible on rejected traffic.
A Reusable withRateLimit Wrapper
Repeating the limit-check-respond code in every handler gets old. Wrap it once:
// lib/rate-limit/with-rate-limit.ts
import type { Ratelimit } from "@upstash/ratelimit";
import { after } from "next/server";
import { getClientIp } from "@/lib/client-ip";
type Handler<Ctx> = (request: Request, context: Ctx) => Promise<Response>;
export function withRateLimit<Ctx>(
limiter: Ratelimit,
handler: Handler<Ctx>,
getKey: (request: Request) => string | Promise<string> = getClientIp,
): Handler<Ctx> {
return async (request, context) => {
const key = await getKey(request);
const { success, limit, remaining, reset, pending } =
await limiter.limit(key);
after(() => pending);
const rateHeaders = {
"RateLimit-Limit": String(limit),
"RateLimit-Remaining": String(remaining),
"RateLimit-Reset": String(
Math.max(0, Math.ceil((reset - Date.now()) / 1000)),
),
};
if (!success) {
return Response.json(
{ error: "Too many requests" },
{
status: 429,
headers: {
...rateHeaders,
"Retry-After": rateHeaders["RateLimit-Reset"],
},
},
);
}
const response = await handler(request, context);
for (const [name, value] of Object.entries(rateHeaders)) {
response.headers.set(name, value);
}
return response;
};
}
The wrapper is generic over the handler's context argument, so it works with dynamic routes whose params is a Promise. It adds quota headers to successful responses and short-circuits with a 429 otherwise.
Using it with a dynamic route and a per-user key:
// app/api/projects/[id]/export/route.ts
import { apiLimiter } from "@/lib/rate-limit/upstash";
import { withRateLimit } from "@/lib/rate-limit/with-rate-limit";
import { getSessionUserId } from "@/lib/auth";
type Ctx = { params: Promise<{ id: string }> };
export const POST = withRateLimit<Ctx>(
apiLimiter,
async (_request, { params }) => {
const { id } = await params;
// ...start the export job
return Response.json({ status: "queued", projectId: id }, { status: 202 });
},
async () => `user:${await getSessionUserId()}`,
);
getSessionUserId stands in for however your app reads the current user (a session cookie, an auth library). Keying by user means a user behind a shared office IP isn't throttled because of their colleagues. If you haven't set up authentication yet, managing authentication in Next.js covers the options.
Rate Limiting in proxy.ts
Route-level limits are precise, but sometimes you want a blanket limit across many routes, for example all of /api, before any handler code runs. That's what proxy.ts (the Next.js 16 name for middleware) is for:
// proxy.ts
import { NextResponse, type NextRequest } from "next/server";
import { Ratelimit } from "@upstash/ratelimit";
import { Redis } from "@upstash/redis";
const globalLimiter = new Ratelimit({
redis: Redis.fromEnv(),
limiter: Ratelimit.slidingWindow(300, "1 m"),
prefix: "rl:global",
});
export async function proxy(request: NextRequest) {
const ip =
request.headers.get("x-forwarded-for")?.split(",")[0].trim() ?? "unknown";
const { success, reset } = await globalLimiter.limit(ip);
if (!success) {
return NextResponse.json(
{ error: "Too many requests" },
{
status: 429,
headers: {
"Retry-After": String(
Math.max(1, Math.ceil((reset - Date.now()) / 1000)),
),
},
},
);
}
return NextResponse.next();
}
export const config = {
matcher: "/api/:path*",
};
The matcher restricts the proxy to API routes, so static assets and pages aren't counted. A good layered setup is a generous global limit in proxy.ts to stop floods, plus tighter per-route limits in sensitive handlers like login, signup, and anything that costs money. For more on what proxy.ts can do, see understanding middleware in Next.js.
Server Actions Need Limits Too
Server Actions are public HTTP endpoints, even though you call them like functions. A contact form action can be spammed exactly like a Route Handler. The same limiter works there; read the IP with headers():
// app/contact/actions.ts
"use server";
import { headers } from "next/headers";
import { Ratelimit } from "@upstash/ratelimit";
import { Redis } from "@upstash/redis";
const contactLimiter = new Ratelimit({
redis: Redis.fromEnv(),
limiter: Ratelimit.fixedWindow(3, "1 h"),
prefix: "rl:contact",
});
export async function sendContactMessage(formData: FormData) {
const headerList = await headers();
const ip =
headerList.get("x-forwarded-for")?.split(",")[0].trim() ?? "unknown";
const { success } = await contactLimiter.limit(ip);
if (!success) {
return {
ok: false,
error: "You've sent too many messages. Please try later.",
};
}
const message = String(formData.get("message") ?? "");
// ...send the message
return { ok: true, length: message.length };
}
Return an error state rather than throwing, so the form can show a friendly message.
Testing Your Limits
A quick way to confirm a limiter works is a loop in the terminal:
for i in $(seq 1 25); do
curl -s -o /dev/null -w "%{http_code} " "http://localhost:3000/api/search?q=test"
done
You should see a run of 200s followed by 429s. Check that Retry-After is sensible with curl -i. In automated tests, inject the limiter (or use a very small limit and a separate Redis prefix) so tests don't interfere with each other.
Where Else to Limit
Application-level rate limiting is one layer. It works best combined with:
- Your host or CDN's firewall (Cloudflare, Vercel Firewall, AWS WAF) for volumetric attacks that you never want reaching Node.js at all.
- A reverse proxy like nginx with
limit_reqwhen self-hosting. - Per-account quotas in your database for business limits ("1,000 API calls per month on the free plan"), which are about billing rather than abuse.
Conclusion
Rate limiting a Next.js Route Handler comes down to three decisions: the algorithm (sliding window is a good default, token bucket for bursty clients), the key (user or API key where possible, IP for anonymous endpoints, read from headers you trust), and where state lives (memory for a single server, Redis for anything distributed or serverless). Return a fast 429 with Retry-After, wrap the logic once so every handler can reuse it, add a broad limit in proxy.ts for floods, and don't forget that Server Actions are endpoints too.


