
Creating sitemap.xml and robots.txt Programmatically in Next.js
A sitemap tells search engines which URLs exist on your site and when they last changed. A robots.txt tells crawlers where they're allowed to go and where to find that sitemap. Neither is glamorous, and both are easy to get subtly wrong: a sitemap full of URLs that redirect, a lastmod that changes on every build, or a preview deployment that gets indexed because its robots.txt said "come on in".
You could write these files by hand and drop them in public/, but on any site with dynamic content they go stale the day you publish something new. The App Router has file conventions for both, app/sitemap.ts and app/robots.ts, that generate them from code using the same data as your pages. This post covers both, plus splitting large sitemaps, keeping them fresh, and the mistakes worth avoiding.
Static Files vs Generated Files
Both files can be static or generated:
| File | Static | Generated |
|---|---|---|
| Sitemap | app/sitemap.xml | app/sitemap.ts |
| Robots | app/robots.txt | app/robots.ts |
The static versions are fine for a five-page marketing site that rarely changes. For anything with a blog, a product catalog, or user-generated pages, generate them. The generated versions are special Route Handlers: they're served at /sitemap.xml and /robots.txt, and they're cached by default, so the code runs at build time unless you use request-time APIs.
Don't put a sitemap.xml or robots.txt in public/ and also have the app/ version. Pick one place.
Your First Generated Sitemap
A sitemap file default-exports a function that returns an array of entries. Next.js turns that into valid XML.
// app/sitemap.ts
import type { MetadataRoute } from "next";
const BASE_URL = "https://tidewave.dev";
export default function sitemap(): MetadataRoute.Sitemap {
return [
{ url: BASE_URL, lastModified: "2026-09-01" },
{ url: `${BASE_URL}/about`, lastModified: "2026-03-14" },
{ url: `${BASE_URL}/blog`, lastModified: "2026-09-30" },
];
}
Visit /sitemap.xml and you'll get:
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://tidewave.dev</loc>
<lastmod>2026-09-01</lastmod>
</url>
<url>
<loc>https://tidewave.dev/about</loc>
<lastmod>2026-03-14</lastmod>
</url>
<url>
<loc>https://tidewave.dev/blog</loc>
<lastmod>2026-09-30</lastmod>
</url>
</urlset>
Each entry supports these fields:
| Field | Type | Notes |
|---|---|---|
url | string | Must be absolute. Required. |
lastModified | string or Date | When the page content last changed. |
changeFrequency | "always" ... "never" | Hint only; Google ignores it. |
priority | number (0 to 1) | Hint only; Google ignores it. |
alternates.languages | object | hreflang alternates for translated pages. |
images | string[] | Image URLs for an image sitemap. |
videos | array | Video metadata for a video sitemap. |
Unlike most metadata fields, sitemap URLs don't use metadataBase, so they have to be absolute. Keep the base URL in one place, ideally an environment variable, so preview and production builds produce correct URLs.
Adding Dynamic Content
The point of generating the sitemap is to include every post, product, or page from your data. The function can be async, so you can read files, query a database, or call a CMS:
// lib/site.ts
export const BASE_URL =
process.env.NEXT_PUBLIC_SITE_URL?.replace(/\/$/, "") ??
"http://localhost:3000";
// app/sitemap.ts
import type { MetadataRoute } from "next";
import { BASE_URL } from "@/lib/site";
import { getAllPosts, getAllTags } from "@/lib/posts";
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const posts = await getAllPosts(); // published posts only
const tags = await getAllTags();
const latestPostDate = posts
.map((p) => p.updatedAt ?? p.publishedAt)
.sort()
.at(-1);
const staticPages: MetadataRoute.Sitemap = [
{ url: BASE_URL, lastModified: latestPostDate },
{ url: `${BASE_URL}/blog`, lastModified: latestPostDate },
{ url: `${BASE_URL}/about`, lastModified: "2026-03-14" },
];
const postPages: MetadataRoute.Sitemap = posts.map((post) => ({
url: `${BASE_URL}/blog/${post.slug}`,
lastModified: post.updatedAt ?? post.publishedAt,
images: post.coverImage ? [`${BASE_URL}${post.coverImage}`] : undefined,
}));
const tagPages: MetadataRoute.Sitemap = tags.map((tag) => ({
url: `${BASE_URL}/tags/${tag.slug}`,
}));
return [...staticPages, ...postPages, ...tagPages];
}
What's happening here:
- Posts come from the same source as the blog pages. If
getAllPostsfilters out drafts and future-dated posts for the blog index, the sitemap automatically excludes them too. This is the main advantage over a hand-written file. lastModifiedreflects real content changes. Each post uses its updated date, falling back to its publish date. The home and blog index pages use the date of the newest post, because that's when their content last changed.- Images are included as an image sitemap extension, which can help image search discover cover images. They must be absolute URLs as well.
- Tag pages have no
lastModified, since there's no single meaningful date. Omitting it is better than inventing one.
If you're using dynamic routes with generateStaticParams, you're usually iterating over the same data there. Sharing one getAllPosts function between generateStaticParams, the page, and the sitemap keeps them consistent. For background on how those routes are built, see dynamic routes in Next.js.
Don't use new Date() for lastModified
The docs' quick examples use lastModified: new Date(), and it's tempting to copy that. Don't. It stamps every URL with the build time, so every deploy tells search engines that every page changed. Google uses lastmod only when it's consistently accurate; if it's always "now", it learns to ignore it. Use real dates from your content, or leave the field out.
What to leave out
A sitemap should list canonical, indexable URLs that return a 200. Leave out:
- Pages marked
noindex(search results, thank-you pages, account pages) - URLs that redirect, including the non-canonical form if you redirect
/blog/to/blog - Paginated archive pages beyond the first, unless they carry unique value
- Drafts, previews, and anything behind authentication
- URLs with tracking or filter query strings
If a page's metadata sets alternates.canonical to a different URL, the sitemap should list that canonical URL, not the variant.
Localized Sitemaps
For multilingual sites, add alternates.languages to each entry. Next.js outputs xhtml:link elements with hreflang attributes:
// app/sitemap.ts
import type { MetadataRoute } from "next";
import { BASE_URL } from "@/lib/site";
const locales = ["en", "de", "es"] as const;
const paths = ["", "/about", "/pricing"];
export default function sitemap(): MetadataRoute.Sitemap {
return paths.map((path) => ({
url: `${BASE_URL}/en${path}`,
alternates: {
languages: Object.fromEntries(
locales.map((locale) => [locale, `${BASE_URL}/${locale}${path}`]),
),
},
}));
}
Each URL lists every language version, including itself. Search engines use this to show the right version to the right audience. The same alternates should appear in each page's metadata, so keep both generated from the same list of locales.
Large Sites: Splitting the Sitemap
A single sitemap file can contain at most 50,000 URLs (and 50 MB uncompressed). Large sites need to split it. There are two ways.
Nested sitemap files
You can put a sitemap.ts in any route segment. app/sitemap.ts becomes /sitemap.xml and app/blog/sitemap.ts becomes /blog/sitemap.xml. This works well when your content naturally falls into sections.
generateSitemaps
For one big collection, export generateSitemaps to produce several sitemaps from one file:
// app/products/sitemap.ts
import type { MetadataRoute } from "next";
import { BASE_URL } from "@/lib/site";
import { countProducts, getProductsPage } from "@/lib/products";
const PER_SITEMAP = 50_000;
export async function generateSitemaps() {
const total = await countProducts();
const count = Math.max(1, Math.ceil(total / PER_SITEMAP));
return Array.from({ length: count }, (_, i) => ({ id: i }));
}
export default async function sitemap(props: {
id: Promise<string>;
}): Promise<MetadataRoute.Sitemap> {
const id = Number(await props.id);
const products = await getProductsPage({
offset: id * PER_SITEMAP,
limit: PER_SITEMAP,
});
return products.map((product) => ({
url: `${BASE_URL}/products/${product.slug}`,
lastModified: product.updatedAt,
}));
}
generateSitemaps returns one object per sitemap, each with an id. The default export receives that id as a promise resolving to a string (a Next.js 16 change), so convert it with Number() before doing arithmetic. The files are served at /products/sitemap/0.xml, /products/sitemap/1.xml, and so on.
Pointing crawlers at multiple sitemaps
Next.js doesn't create a sitemap index file for you. The simplest approach is to list every sitemap in robots.txt, which accepts more than one Sitemap: line (shown in the next section). If you prefer a single index URL to submit in search consoles, a small Route Handler can produce one:
// app/sitemap-index.xml/route.ts
import { BASE_URL } from "@/lib/site";
import { generateSitemaps } from "@/app/products/sitemap";
export async function GET() {
const productSitemaps = await generateSitemaps();
const urls = [
`${BASE_URL}/sitemap.xml`,
...productSitemaps.map(
({ id }) => `${BASE_URL}/products/sitemap/${id}.xml`,
),
];
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
${urls.map((loc) => ` <sitemap><loc>${loc}</loc></sitemap>`).join("\n")}
</sitemapindex>`;
return new Response(xml, {
headers: { "Content-Type": "application/xml" },
});
}
The folder name sitemap-index.xml becomes the URL path, so this is served at /sitemap-index.xml. It reuses generateSitemaps so the index always matches the files that exist.
Generating robots.txt
app/robots.ts default-exports a function returning a Robots object:
// app/robots.ts
import type { MetadataRoute } from "next";
import { BASE_URL } from "@/lib/site";
export default function robots(): MetadataRoute.Robots {
return {
rules: {
userAgent: "*",
allow: "/",
disallow: ["/api/", "/account/", "/search"],
},
sitemap: `${BASE_URL}/sitemap.xml`,
};
}
Output at /robots.txt:
User-Agent: *
Allow: /
Disallow: /api/
Disallow: /account/
Disallow: /search
Sitemap: https://tidewave.dev/sitemap.xml
sitemap accepts an array, which is how you point crawlers at several sitemaps:
sitemap: [
`${BASE_URL}/sitemap.xml`,
`${BASE_URL}/products/sitemap/0.xml`,
`${BASE_URL}/products/sitemap/1.xml`,
],
Rules for specific crawlers
Pass an array to rules to give different crawlers different instructions. A common use is opting out of crawlers used to collect AI training data while staying open to search engines:
// app/robots.ts
import type { MetadataRoute } from "next";
import { BASE_URL } from "@/lib/site";
export default function robots(): MetadataRoute.Robots {
return {
rules: [
{ userAgent: "*", allow: "/", disallow: ["/api/", "/account/"] },
{ userAgent: ["GPTBot", "CCBot", "Google-Extended"], disallow: "/" },
],
sitemap: `${BASE_URL}/sitemap.xml`,
};
}
Each object becomes its own User-Agent block. Check each crawler operator's documentation for the exact user agent token, since they change from time to time. robots.txt is a request, not access control: well-behaved crawlers follow it, others don't.
Since Next.js 16.3, rules also accept an other field for non-standard directives, like Request-Rate for specific engines. Values are passed through verbatim.
Blocking non-production deployments
Preview and staging deployments should never be indexed. Because robots.ts is code, you can switch on an environment variable:
// app/robots.ts
import type { MetadataRoute } from "next";
import { BASE_URL } from "@/lib/site";
const isProduction = process.env.SITE_ENV === "production";
export default function robots(): MetadataRoute.Robots {
if (!isProduction) {
return { rules: { userAgent: "*", disallow: "/" } };
}
return {
rules: { userAgent: "*", allow: "/", disallow: ["/api/", "/account/"] },
sitemap: `${BASE_URL}/sitemap.xml`,
};
}
Set SITE_ENV=production only in your production environment (or use whatever variable your host provides for this). Pair it with a noindex robots meta tag in the root layout for non-production builds, because a Disallow stops crawling but doesn't remove URLs a search engine already knows about. The Metadata API post shows that pattern.
Disallow is not noindex
This catches a lot of people. Disallow: /search stops crawlers from fetching /search, but if other sites link to it, the URL can still appear in results (without a description). To keep a page out of search results, let it be crawled and give it a noindex robots meta tag. Use Disallow to save crawl budget on things like API routes and infinite filter combinations, not as a privacy tool.
Keeping the Sitemap Fresh
Because sitemap.ts is cached by default, it's generated at build time. For a site that rebuilds on every content change (like a Markdown blog deployed from Git), that's exactly right.
If content changes without a rebuild, for example posts in a CMS or database, you need the sitemap to refresh. How depends on whether you've enabled Cache Components.
Without Cache Components, use the revalidate route segment option to regenerate the sitemap periodically:
// app/sitemap.ts
export const revalidate = 3600; // regenerate at most once an hour
With Cache Components (cacheComponents: true), cache the data function instead:
// lib/posts.ts
import { cacheLife, cacheTag } from "next/cache";
export async function getAllPosts() {
"use cache";
cacheLife("hours");
cacheTag("posts");
const res = await fetch("https://cms.example.com/api/posts?status=published");
return (await res.json()) as {
slug: string;
publishedAt: string;
updatedAt?: string;
coverImage?: string;
}[];
}
Either way, when your CMS sends a publish webhook, you can call revalidateTag("posts", "max") (or revalidatePath("/sitemap.xml")) to refresh immediately rather than waiting. Our post on on-demand revalidation covers wiring that up.
Common Problems
The sitemap URL is blocked by Proxy. If you have a proxy.ts that redirects unauthenticated users, make sure its matcher excludes /sitemap.xml, /robots.txt, and any nested sitemap paths. A matcher with a negative lookahead like "/((?!api|_next/static|_next/image|favicon.ico|sitemap.xml|robots.txt).*)" is the usual approach.
URLs use localhost in production. The base URL wasn't set at build time. Since the sitemap is built during next build, the environment variable has to be available then, not only at runtime.
Trailing slash mismatch. If trailingSlash: true is set in next.config.ts, every URL in the sitemap should end with /. Otherwise you're listing URLs that redirect.
Drafts in the sitemap. Make sure the sitemap reads from the same filtered function as your public pages, not directly from the raw data source.
Validation errors. Open /sitemap.xml in the browser and check it renders as XML. After deploying, submit it in Google Search Console and Bing Webmaster Tools and check the reported errors, which point out non-200 URLs and malformed entries.
Conclusion
app/sitemap.ts and app/robots.ts turn two easily neglected files into code that stays in sync with your content. Build the sitemap from the same data functions your pages use, give each URL a real lastModified (or none), leave out anything non-canonical or noindex, and split with generateSitemaps once you approach 50,000 URLs. Use robots.ts to point at your sitemaps, keep crawlers out of API and account routes, and block every non-production deployment. Then check both files in Search Console once after launch, and again any time you change your URL structure.


