Type something to search...
Streaming Responses and Server-Sent Events with Next.js Route Handlers

Streaming Responses and Server-Sent Events with Next.js Route Handlers

Most HTTP responses are all-or-nothing: the server does its work, then sends the whole body at once. That's fine for a list of products. It's not fine when the work takes thirty seconds, when the answer arrives token by token from a language model, or when the server needs to tell the browser about something that just happened. For those cases you want to start sending before you're finished.

Next.js Route Handlers are built on the Web Response API, and a Response body can be a ReadableStream. That one fact gets you chunked streaming, progress updates, and server-sent events (SSE) without any extra libraries. This post covers how to stream plain text from a Route Handler, how to read a stream in the browser, how the SSE format works, how to build an SSE endpoint that cleans up after itself, how to stream from POST requests, and the infrastructure details (buffering, compression, timeouts) that make streaming silently stop working in production.

This post is about streaming from your own HTTP endpoints. If you're after streaming React UI with loading.tsx and Suspense, see loading.tsx and streaming UI instead.

Streaming a Plain Response

Here's the smallest useful streaming handler. It sends a line every half second, five times, then closes:

// app/api/countdown/route.ts
const encoder = new TextEncoder();

function sleep(ms: number) {
  return new Promise((resolve) => setTimeout(resolve, ms));
}

export async function GET() {
  const stream = new ReadableStream<Uint8Array>({
    async start(controller) {
      for (let i = 5; i > 0; i--) {
        controller.enqueue(encoder.encode(`${i}...\n`));
        await sleep(500);
      }
      controller.enqueue(encoder.encode("Liftoff!\n"));
      controller.close();
    },
  });

  return new Response(stream, {
    headers: {
      "Content-Type": "text/plain; charset=utf-8",
      "Cache-Control": "no-cache",
    },
  });
}

The pieces:

  • ReadableStream is the standard Web Streams API, available globally in the Node.js runtime Next.js uses.
  • start(controller) runs once when the stream is created. You push data with controller.enqueue() and finish with controller.close().
  • Streams carry bytes, so strings go through a TextEncoder first.
  • Returning new Response(stream) hands the stream to Next.js, which sends each chunk as it's enqueued.

Notice the handler returns the Response immediately; the loop inside start keeps running after that. That's the whole trick: the response headers go out right away and the body follows over time.

Generators Make This Nicer

If your data comes from an async loop, an async generator is a more natural way to write it. A small helper turns any async iterator into a stream, pulling the next value only when the consumer is ready for it:

// lib/iterator-to-stream.ts
export function iteratorToStream(
  iterator: AsyncIterator<string>,
): ReadableStream<Uint8Array> {
  const encoder = new TextEncoder();
  return new ReadableStream({
    async pull(controller) {
      const { value, done } = await iterator.next();
      if (done) {
        controller.close();
      } else {
        controller.enqueue(encoder.encode(value));
      }
    },
    async cancel() {
      await iterator.return?.();
    },
  });
}
// app/api/report/route.ts
import { iteratorToStream } from "@/lib/iterator-to-stream";

async function* buildReport() {
  yield "Collecting orders...\n";
  await new Promise((r) => setTimeout(r, 800));
  yield "Aggregating totals...\n";
  await new Promise((r) => setTimeout(r, 800));
  yield "Done. Revenue: $12,480\n";
}

export async function GET() {
  return new Response(iteratorToStream(buildReport()), {
    headers: { "Content-Type": "text/plain; charset=utf-8" },
  });
}

Using pull instead of start gives you backpressure: if the client reads slowly, the generator isn't asked for more values until the stream has room. The cancel method runs when the client disconnects, and calling iterator.return() stops the generator so it doesn't keep doing work for nobody.

Reading a Stream in the Browser

fetch exposes the body as a stream too. Instead of await res.text(), which waits for everything, read it chunk by chunk:

// app/report/report-viewer.tsx
"use client";

import { useState } from "react";

export function ReportViewer() {
  const [output, setOutput] = useState("");
  const [running, setRunning] = useState(false);

  async function run() {
    setOutput("");
    setRunning(true);

    const res = await fetch("/api/report");
    if (!res.ok || !res.body) {
      setOutput("Request failed");
      setRunning(false);
      return;
    }

    const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
    while (true) {
      const { value, done } = await reader.read();
      if (done) break;
      setOutput((prev) => prev + value);
    }
    setRunning(false);
  }

  return (
    <div>
      <button type="button" onClick={run} disabled={running}>
        {running ? "Running…" : "Generate report"}
      </button>
      <pre aria-live="polite">{output}</pre>
    </div>
  );
}

TextDecoderStream converts bytes back to text and correctly handles multi-byte characters that get split across chunks (an emoji or accented letter can straddle two chunks). Each reader.read() resolves when the next chunk arrives, so the pre fills in progressively.

One thing to keep in mind: chunk boundaries are not message boundaries. Proxies and the network can merge or split chunks, so if you send structured data you need a framing format. That's exactly what server-sent events give you.

Server-Sent Events: The Format

SSE is a tiny text protocol on top of a streaming HTTP response with Content-Type: text/event-stream. Each event is a few field: value lines followed by a blank line:

id: 42
event: progress
data: {"percent":60}

: this line is a comment, often used as a heartbeat

data: a message with no event name

The fields:

FieldMeaning
dataThe payload. Multiple data lines in one event are joined with newlines.
eventAn event name. Without it, the browser fires a generic message event.
idAn event ID. The browser remembers the last one and sends it back on reconnect.
retryHow many milliseconds the browser should wait before reconnecting.
Lines starting with :Comments. Ignored by the client, useful to keep the connection alive.

The browser side is the built-in EventSource API, which parses this format, dispatches events, and reconnects automatically if the connection drops. Compared with WebSockets, SSE is one-directional (server to client), but it's plain HTTP: it works through most proxies, uses your existing cookies for auth, and needs no special server. And unlike WebSockets, it works with Route Handlers, which can't hold a WebSocket connection.

Building an SSE Endpoint

Here's a reusable helper that wraps the formatting and cleanup, followed by an endpoint that uses it.

// lib/sse.ts
const encoder = new TextEncoder();

type SendFn = (event: { data: unknown; event?: string; id?: string }) => void;

export function createSSEStream(
  request: Request,
  onStart: (send: SendFn) => (() => void) | void,
) {
  let cleanup: (() => void) | void;
  let heartbeat: ReturnType<typeof setInterval> | undefined;
  let closed = false;

  const stop = () => {
    if (closed) return;
    closed = true;
    clearInterval(heartbeat);
    cleanup?.();
  };

  const stream = new ReadableStream<Uint8Array>({
    start(controller) {
      const write = (text: string) => {
        if (!closed) controller.enqueue(encoder.encode(text));
      };

      const send: SendFn = ({ data, event, id }) => {
        let message = "";
        if (id) message += `id: ${id}\n`;
        if (event) message += `event: ${event}\n`;
        message += `data: ${JSON.stringify(data)}\n\n`;
        write(message);
      };

      write("retry: 5000\n\n");
      heartbeat = setInterval(() => write(": ping\n\n"), 15_000);
      cleanup = onStart(send);

      request.signal.addEventListener("abort", () => {
        stop();
        try {
          controller.close();
        } catch {
          // already closed
        }
      });
    },
    cancel() {
      stop();
    },
  });

  return new Response(stream, {
    headers: {
      "Content-Type": "text/event-stream; charset=utf-8",
      "Cache-Control": "no-cache, no-transform",
      "X-Accel-Buffering": "no",
    },
  });
}

What each part does:

  • send formats one event. JSON.stringify guarantees the payload has no raw newlines, which would otherwise break the data: line framing.
  • retry: 5000 tells the browser to wait five seconds before reconnecting after a drop, instead of the browser default (typically around three seconds).
  • The heartbeat sends a comment every 15 seconds. Proxies and load balancers often kill connections that are idle for 30 to 60 seconds; a comment keeps bytes flowing without triggering any event on the client.
  • Cleanup runs from two places: request.signal fires abort when the client disconnects, and the stream's cancel() runs when the consumer cancels the stream. Either way, stop() clears the timers and calls whatever cleanup the caller returned. Without this, every closed browser tab would leave an interval running on your server forever.
  • Headers: text/event-stream is required for EventSource. no-transform asks intermediaries not to compress or rewrite the stream, and X-Accel-Buffering: no tells Nginx not to buffer it (more on that below).

Now an endpoint that streams the server time every second:

// app/api/clock/route.ts
import { createSSEStream } from "@/lib/sse";

export async function GET(request: Request) {
  return createSSEStream(request, (send) => {
    let count = 0;
    const tick = () =>
      send({
        event: "time",
        id: String(++count),
        data: { now: new Date().toISOString() },
      });

    tick();
    const interval = setInterval(tick, 1000);
    return () => clearInterval(interval);
  });
}

The callback returns its own cleanup function, which the helper calls on disconnect. Because the handler reads from request (its signal), it always runs per request and is never prerendered or cached. If you write a streaming handler that doesn't touch the request at all and you have Cache Components enabled, call await connection() from next/server at the top to make sure it isn't prerendered at build time.

Consuming SSE with EventSource

On the client, EventSource handles parsing and reconnection. Wrap it in an effect and close it on unmount:

// app/clock/live-clock.tsx
"use client";

import { useEffect, useState } from "react";

type Status = "connecting" | "open" | "reconnecting";

export function LiveClock() {
  const [now, setNow] = useState<string | null>(null);
  const [status, setStatus] = useState<Status>("connecting");

  useEffect(() => {
    const source = new EventSource("/api/clock");

    source.onopen = () => setStatus("open");
    source.onerror = () => setStatus("reconnecting");

    source.addEventListener("time", (event: MessageEvent<string>) => {
      const payload = JSON.parse(event.data) as { now: string };
      setNow(payload.now);
    });

    return () => source.close();
  }, []);

  return (
    <p>
      {now ? new Date(now).toLocaleTimeString() : "…"} <small>({status})</small>
    </p>
  );
}

Named events (event: time on the server) are received with addEventListener("time", ...). Unnamed events go to onmessage. The onerror handler fires when the connection drops; EventSource then reconnects on its own, so you usually just update the UI rather than doing anything yourself.

The return () => source.close() matters. Without it, every client-side navigation away from this page leaves a connection open, and browsers cap HTTP/1.1 connections at about six per domain. Leak a few and the rest of your app's requests start hanging. Over HTTP/2 the limit is much higher, but closing connections you no longer need is still the right thing to do.

Finite Streams Need an Explicit Close

Here's a surprise that catches most people: when the server closes an SSE stream normally, EventSource treats it as a dropped connection and reconnects. For an endless feed that's what you want. For a finite job, it means the job restarts.

The fix is to send a final event and have the client close the connection itself:

// app/api/jobs/[id]/progress/route.ts
import { createSSEStream } from "@/lib/sse";

export async function GET(
  request: Request,
  { params }: { params: Promise<{ id: string }> },
) {
  const { id } = await params;

  return createSSEStream(request, (send) => {
    let percent = 0;
    const interval = setInterval(() => {
      percent = Math.min(percent + 10, 100);
      send({ event: "progress", data: { jobId: id, percent } });
      if (percent === 100) {
        send({ event: "done", data: { jobId: id } });
        clearInterval(interval);
      }
    }, 500);
    return () => clearInterval(interval);
  });
}
// app/jobs/[id]/job-progress.tsx
"use client";

import { useEffect, useState } from "react";

export function JobProgress({ jobId }: { jobId: string }) {
  const [percent, setPercent] = useState(0);
  const [done, setDone] = useState(false);

  useEffect(() => {
    const source = new EventSource(`/api/jobs/${jobId}/progress`);

    source.addEventListener("progress", (event: MessageEvent<string>) => {
      setPercent((JSON.parse(event.data) as { percent: number }).percent);
    });

    source.addEventListener("done", () => {
      setDone(true);
      source.close(); // stop the automatic reconnect
    });

    return () => source.close();
  }, [jobId]);

  return (
    <div>
      <progress value={percent} max={100} />
      <p>{done ? "Finished" : `${percent}%`}</p>
    </div>
  );
}

The simulated job here uses an interval. In a real app the progress would come from wherever the work happens, which leads to the next point.

Where Do Events Come From?

In the examples so far the handler generates its own events. Real apps usually need to push events that originate elsewhere: a webhook arrives, a background worker finishes, another user posts a comment. Inside a single long-running Node.js server, an in-process EventEmitter can connect the two. On serverless or multi-instance deployments, the request that triggers the event and the request holding the SSE connection usually run in different processes, so you need a shared channel such as Redis pub/sub, Postgres LISTEN/NOTIFY, or a hosted real-time service. The SSE handler subscribes when the connection opens and unsubscribes in its cleanup function, which is exactly what the helper's cleanup hook is for.

Resuming After a Reconnect

When EventSource reconnects, it sends a Last-Event-ID header with the last id it received. If your events are stored (say, rows with incrementing IDs), you can replay what the client missed:

// app/api/notifications/route.ts
import { createSSEStream } from "@/lib/sse";

type Notification = { id: number; text: string };

// Replace with a real query: SELECT ... WHERE id > $1 ORDER BY id
async function getNotificationsSince(lastId: number): Promise<Notification[]> {
  const all: Notification[] = [
    { id: 1, text: "Welcome!" },
    { id: 2, text: "Your report is ready" },
  ];
  return all.filter((n) => n.id > lastId);
}

export async function GET(request: Request) {
  const lastId = Number(request.headers.get("last-event-id") ?? 0) || 0;
  const missed = await getNotificationsSince(lastId);

  return createSSEStream(request, (send) => {
    for (const n of missed) {
      send({ event: "notification", id: String(n.id), data: n });
    }
    // ...then subscribe to new notifications and return an unsubscribe function
  });
}

Streaming from a POST Request

EventSource only does GET requests and can't set custom headers or send a body. For things like "send this prompt and stream the answer", use fetch with a POST and parse the SSE frames yourself. The server side is the same helper:

// app/api/generate/route.ts
import { createSSEStream } from "@/lib/sse";

export async function POST(request: Request) {
  const { prompt } = (await request.json()) as { prompt: string };
  const words = `You asked: ${prompt}. Here is a streamed answer.`.split(" ");

  return createSSEStream(request, (send) => {
    let i = 0;
    const interval = setInterval(() => {
      if (i < words.length) {
        send({ event: "token", data: words[i++] + " " });
      } else {
        send({ event: "done", data: null });
        clearInterval(interval);
      }
    }, 120);
    return () => clearInterval(interval);
  });
}

On the client, buffer incoming text and split on the blank line that ends each event:

// lib/read-sse.ts
export async function readSSE(
  response: Response,
  onEvent: (event: { event: string; data: string }) => void,
) {
  if (!response.body) throw new Error("No response body");

  const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
  let buffer = "";

  while (true) {
    const { value, done } = await reader.read();
    if (done) break;
    buffer += value;

    const frames = buffer.split("\n\n");
    buffer = frames.pop() ?? "";

    for (const frame of frames) {
      let event = "message";
      const data: string[] = [];
      for (const line of frame.split("\n")) {
        if (line.startsWith("event:")) event = line.slice(6).trim();
        else if (line.startsWith("data:")) data.push(line.slice(5).trimStart());
      }
      if (data.length > 0) onEvent({ event, data: data.join("\n") });
    }
  }
}

The last element of frames is whatever arrived after the final blank line, possibly half an event, so it goes back into the buffer to be completed by the next chunk. Usage from a Client Component:

const res = await fetch("/api/generate", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ prompt }),
});

await readSSE(res, ({ event, data }) => {
  if (event === "token") setAnswer((prev) => prev + JSON.parse(data));
});

Pass an AbortController's signal to fetch if users can cancel; aborting on the client fires request.signal on the server, which runs the cleanup.

If you're streaming from a language model, you rarely need to write this by hand. Provider SDKs and the Vercel AI SDK return ready-made streaming Response objects. See building an AI chat interface with the Vercel AI SDK for that path.

Proxying an Upstream Stream

If an upstream API already streams, you don't have to read and re-emit it. Pass its body straight through:

// app/api/upstream/route.ts
export async function GET(request: Request) {
  const upstream = await fetch("https://api.example.com/events", {
    headers: { Authorization: `Bearer ${process.env.UPSTREAM_TOKEN}` },
    signal: request.signal,
  });

  if (!upstream.ok || !upstream.body) {
    return new Response("Upstream error", { status: 502 });
  }

  return new Response(upstream.body, {
    headers: {
      "Content-Type": "text/event-stream; charset=utf-8",
      "Cache-Control": "no-cache, no-transform",
    },
  });
}

This keeps the secret token on the server while the browser receives the stream. Forwarding request.signal means a client disconnect also cancels the upstream request.

Why Streaming Breaks in Production

Streaming that works perfectly with next dev often arrives in one lump once deployed. The culprit is almost always something between your handler and the browser collecting the whole response before passing it on.

  • Reverse proxies. Nginx buffers responses by default. The X-Accel-Buffering: no header in the helper turns that off per response. Other proxies have their own settings.
  • Compression. Gzip and Brotli encoders buffer data until they have enough to compress efficiently. Cache-Control: no-transform discourages intermediaries from compressing; if your own server layer compresses, make sure it flushes per chunk or skips text/event-stream.
  • CDNs and load balancers. Some buffer by default or only pass streams through on certain plans or configurations. Check your provider's docs.
  • Serverless timeouts. Functions have a maximum duration. An SSE connection that outlives it gets cut, and EventSource reconnects (which is why retry and Last-Event-ID are worth setting up). Where your host supports it, raise the limit per route with export const maxDuration = 300.
  • Runtime. Streaming works on the default Node.js runtime. The Edge runtime is deprecated in Next.js 16, so don't add export const runtime = "edge" for streaming.
  • Testing with curl. curl buffers output too. Use curl -N http://localhost:3000/api/clock to see events as they arrive.

A quick production check: open DevTools, select the request in the Network tab, and look at the EventStream tab (Chrome shows it for text/event-stream responses). Events should appear one at a time, not all together.

When to Use What

NeedApproach
Long text output (reports, LLM answers)Stream the response body, read with fetch
Server pushes updates to the browserSSE with EventSource
Progress of a long jobSSE with a final done event and client-side close()
Streaming answer to a POSTSSE frames over fetch, parsed manually
Two-way, low-latency messaging (chat presence, games)WebSockets via a separate service
Updates every minute or soPlain polling is simpler

Conclusion

A Route Handler can return a ReadableStream, and that's all you need to stream. Use pull and async generators for backpressure-friendly text streams, read them on the client with getReader() and TextDecoderStream, and switch to the SSE format when you need message framing, named events, and automatic reconnection. Always clean up on disconnect via request.signal and cancel(), send heartbeats, close finite streams from the client, and remember that buffering in proxies, compression, and CDNs is the usual reason streaming "doesn't work" after deployment.

Tags :
Share :

Related Posts

A Deep Dive into next.config Options Every Developer Should Know

A Deep Dive into next.config Options Every Developer Should Know

next.config.ts is the one file every Next.js project has and almost nobody reads end to end. It starts as an empty object, then slowly collects a r

Continue Reading
Adding JSON-LD Structured Data to Next.js Pages for Rich Search Results

Adding JSON-LD Structured Data to Next.js Pages for Rich Search Results

Search engines are good at reading pages, but they still guess. Is "4.7" a rating or a version number? Is that date when the article was published or

Continue Reading
Adding Page Transitions and Animations to Next.js with Framer Motion

Adding Page Transitions and Animations to Next.js with Framer Motion

Animation is one of the easiest ways to make an app feel polished, and one of the easiest ways to make it feel slow. A subtle fade when a page loads,

Continue Reading