
Building a Chat UI with React and AI Streaming Responses
A language model can take ten or twenty seconds to write a long answer. If your chat UI waits for the whole response before showing anything, users stare at a spinner and assume it's broken. Every good AI chat interface streams instead: text appears as it's generated, a word or a few words at a time, and the user can start reading almost immediately.
Streaming changes how you build the UI. You're no longer handling one request and one response. You're reading a stream of chunks, appending them to a message that's still being written, keeping the scroll position sensible, letting users stop generation halfway, and dealing with errors that can happen in the middle of a reply.
In this post you'll build a complete streaming chat in React 19 with TypeScript. The backend is a small Express endpoint that calls Claude through the Anthropic SDK and forwards the stream. The frontend reads it with fetch and a ReadableStream, manages state in a useChat hook, renders Markdown, and handles stopping, errors, and scrolling.
How the Pieces Fit Together
The browser should never call the model provider directly. Your API key would be visible to anyone who opens DevTools. Instead, the flow looks like this:
- The React app POSTs the conversation to your own endpoint,
/api/chat. - The server calls the model with streaming enabled and the API key from its environment.
- As the model produces text, the server writes each piece to the HTTP response without closing it.
- The browser reads the response body incrementally and updates the UI with each piece.
For the wire format between your server and your client, use something simple you control. Here we'll use newline-delimited JSON (NDJSON): each line is one JSON event, such as a text chunk, a completion marker, or an error. It's easy to produce and parse.
The Server Endpoint
Install the dependencies:
npm install express @anthropic-ai/sdk
npm install -D @types/express tsx
Set ANTHROPIC_API_KEY in the server's environment. The SDK reads it automatically.
// server/index.ts
import express from "express";
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const app = express();
app.use(express.json({ limit: "200kb" }));
type ChatEvent =
| { type: "text"; text: string }
| { type: "done"; stopReason: string | null }
| { type: "error"; message: string };
function isValidHistory(value: unknown): value is Anthropic.MessageParam[] {
return (
Array.isArray(value) &&
value.length > 0 &&
value.length <= 50 &&
value.every(
(m) =>
(m?.role === "user" || m?.role === "assistant") &&
typeof m.content === "string" &&
m.content.length <= 20_000,
) &&
value[0].role === "user"
);
}
app.post("/api/chat", async (req, res) => {
const history: unknown = req.body?.messages;
if (!isValidHistory(history)) {
res.status(400).json({ error: "Invalid messages" });
return;
}
// Stop calling the model if the browser disconnects or the user hits Stop
const controller = new AbortController();
res.on("close", () => controller.abort());
res.setHeader("Content-Type", "application/x-ndjson; charset=utf-8");
res.setHeader("Cache-Control", "no-cache, no-transform");
res.setHeader("X-Accel-Buffering", "no");
const send = (event: ChatEvent) => res.write(JSON.stringify(event) + "\n");
try {
const stream = client.messages.stream(
{
model: "claude-opus-5-5",
max_tokens: 8000,
output_config: { effort: "low" },
system: "You are a concise, friendly assistant for a web development blog. Use Markdown.",
messages: history,
},
{ signal: controller.signal },
);
for await (const event of stream) {
if (event.type === "content_block_delta" && event.delta.type === "text_delta") {
send({ type: "text", text: event.delta.text });
}
}
const final = await stream.finalMessage();
send({ type: "done", stopReason: final.stop_reason });
res.end();
} catch (error) {
if (controller.signal.aborted) return;
console.error(error);
const message =
error instanceof Anthropic.RateLimitError
? "The assistant is busy. Please try again in a moment."
: "Something went wrong while generating a reply.";
if (!res.headersSent) {
res.status(502).json({ error: message });
} else {
send({ type: "error", message });
res.end();
}
}
});
app.listen(3001, () => console.log("API on http://localhost:3001"));
A few details here matter more than they look:
- Validation. The endpoint is public, so it accepts only a short, size-capped list of plain-text messages.
- Aborting. When the user clicks Stop, the response emits
close, and aborting the SDK stream stops generation so you don't pay for unread tokens. - Buffering headers. Proxies such as Nginx may buffer responses and deliver them all at once, which silently defeats streaming.
X-Accel-Buffering: noandno-transformtell them not to. - Errors mid-stream. Once you've started writing, you can't change the status code, so later errors are sent as an
errorevent. - Stop reason. The final
doneevent reports why generation ended. A value ofmax_tokensmeans the reply was cut off, andrefusalmeans the model declined, so the UI can show a note instead of an answer that just stops.
Run it with npx tsx server/index.ts. In a Vite app, proxy /api to it during development so the browser sees a single origin:
// vite.config.ts
import { defineConfig } from "vite";
import react from "@vitejs/plugin-react";
export default defineConfig({
plugins: [react()],
server: {
proxy: { "/api": "http://localhost:3001" },
},
});
Reading a Stream in the Browser
fetch gives you the response body as a ReadableStream of bytes. To turn it into events, decode the bytes to text, split on newlines, and parse each complete line. The one subtlety is that chunks don't align with lines: a network chunk can end halfway through a JSON object, so you need to keep the incomplete tail in a buffer until the rest arrives.
// src/chat/readNdjson.ts
export async function* readNdjson<T>(response: Response): AsyncGenerator<T> {
if (!response.body) throw new Error("Response has no body");
const reader = response.body.pipeThrough(new TextDecoderStream()).getReader();
let buffer = "";
try {
while (true) {
const { value, done } = await reader.read();
if (done) break;
buffer += value;
const lines = buffer.split("\n");
buffer = lines.pop() ?? "";
for (const line of lines) {
if (line.trim()) yield JSON.parse(line) as T;
}
}
if (buffer.trim()) yield JSON.parse(buffer) as T;
} finally {
reader.releaseLock();
}
}
TextDecoderStream handles multi-byte characters split across chunks, which a naive new TextDecoder().decode(chunk) per chunk would garble. As an async generator, it stays separate from React and easy to test.
Modeling Chat State
Each message needs an ID for stable keys and updates, a role, its content, and a status so the UI knows whether it's still streaming:
// src/chat/types.ts
export type Role = "user" | "assistant";
export type ChatMessage = {
id: string;
role: Role;
content: string;
status: "done" | "streaming" | "error" | "stopped";
note?: string;
};
export type ServerEvent =
| { type: "text"; text: string }
| { type: "done"; stopReason: string | null }
| { type: "error"; message: string };
The useChat Hook
The hook owns the message list, sends requests, consumes the stream, and exposes send and stop. Putting this in a custom hook keeps the components purely about presentation, which the post on building your own custom hooks explains in more depth.
// src/chat/useChat.ts
import { useCallback, useRef, useState } from "react";
import { readNdjson } from "./readNdjson";
import type { ChatMessage, ServerEvent } from "./types";
export function useChat() {
const [messages, setMessages] = useState<ChatMessage[]>([]);
const [isStreaming, setIsStreaming] = useState(false);
const abortRef = useRef<AbortController | null>(null);
const updateMessage = useCallback((id: string, update: (m: ChatMessage) => ChatMessage) => {
setMessages((prev) => prev.map((m) => (m.id === id ? update(m) : m)));
}, []);
const send = useCallback(
async (text: string) => {
const content = text.trim();
if (!content || isStreaming) return;
const userMessage: ChatMessage = {
id: crypto.randomUUID(),
role: "user",
content,
status: "done",
};
const assistantId = crypto.randomUUID();
// Only send completed turns as history
const history = [...messages, userMessage]
.filter((m) => m.status === "done" && m.content.length > 0)
.map(({ role, content }) => ({ role, content }));
setMessages((prev) => [
...prev,
userMessage,
{ id: assistantId, role: "assistant", content: "", status: "streaming" },
]);
setIsStreaming(true);
const controller = new AbortController();
abortRef.current = controller;
try {
const res = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages: history }),
signal: controller.signal,
});
if (!res.ok) {
const body = (await res.json().catch(() => null)) as { error?: string } | null;
throw new Error(body?.error ?? `Request failed with ${res.status}`);
}
for await (const event of readNdjson<ServerEvent>(res)) {
if (event.type === "text") {
updateMessage(assistantId, (m) => ({ ...m, content: m.content + event.text }));
} else if (event.type === "done") {
updateMessage(assistantId, (m) => ({
...m,
status: "done",
note:
event.stopReason === "max_tokens"
? "The reply was cut off because it reached the length limit."
: event.stopReason === "refusal"
? "The assistant declined to answer this request."
: undefined,
}));
} else if (event.type === "error") {
throw new Error(event.message);
}
}
} catch (error) {
if (controller.signal.aborted) {
updateMessage(assistantId, (m) => ({ ...m, status: "stopped" }));
} else {
const message = error instanceof Error ? error.message : "Something went wrong.";
updateMessage(assistantId, (m) => ({ ...m, status: "error", note: message }));
}
} finally {
abortRef.current = null;
setIsStreaming(false);
}
},
[messages, isStreaming, updateMessage],
);
const stop = useCallback(() => {
abortRef.current?.abort();
}, []);
return { messages, isStreaming, send, stop };
}
The key pattern is the functional state update in updateMessage. The stream loop holds an old messages snapshot, so setMessages(messages.map(...)) would overwrite earlier chunks with stale data. The functional form always works on the latest state.
Also notice what goes into history. Stopped and errored replies are excluded, so a half-written answer doesn't confuse the next turn. If you'd rather keep partial answers as context, include stopped messages too. Both are reasonable product decisions.
Rendering Messages and Markdown
Models answer in Markdown, with lists, code blocks, and bold text. Render it with react-markdown, which builds React elements rather than injecting HTML, so it's safe by default:
npm install react-markdown remark-gfm
// src/chat/MessageBubble.tsx
import { memo } from "react";
import Markdown from "react-markdown";
import remarkGfm from "remark-gfm";
import type { ChatMessage } from "./types";
export const MessageBubble = memo(function MessageBubble({ message }: { message: ChatMessage }) {
const isUser = message.role === "user";
return (
<div className={`flex ${isUser ? "justify-end" : "justify-start"}`}>
<div
className={`max-w-[80%] rounded-2xl px-4 py-2 ${
isUser ? "bg-blue-600 text-white" : "bg-slate-100 text-slate-900"
}`}
>
{isUser ? (
<p className="whitespace-pre-wrap">{message.content}</p>
) : (
<div className="prose prose-sm max-w-none">
{message.content ? (
<Markdown remarkPlugins={[remarkGfm]}>{message.content}</Markdown>
) : (
<span className="animate-pulse text-slate-500">Thinking...</span>
)}
</div>
)}
{message.status === "stopped" && <p className="mt-1 text-xs text-slate-500">Stopped</p>}
{message.note && (
<p className={`mt-1 text-xs ${message.status === "error" ? "text-red-600" : "text-slate-500"}`}>
{message.note}
</p>
)}
</div>
</div>
);
});
Wrapping the bubble in memo matters for long conversations. Every chunk creates a new messages array, but only the streaming message object changes. The others keep the same reference, so memo skips re-rendering and re-parsing their Markdown. See preventing unnecessary re-renders with React.memo for more on when this pays off.
User messages render as plain text, since there's no reason to interpret Markdown the user typed.
Scrolling That Doesn't Fight the User
A chat should follow new content to the bottom, but only if the user is already at the bottom. If they've scrolled up to reread something, yanking them down on every chunk is infuriating.
// src/chat/useStickToBottom.ts
import { useEffect, useRef } from "react";
export function useStickToBottom<T extends HTMLElement>(dependency: unknown) {
const ref = useRef<T>(null);
const pinned = useRef(true);
useEffect(() => {
const el = ref.current;
if (!el) return;
const onScroll = () => {
const distance = el.scrollHeight - el.scrollTop - el.clientHeight;
pinned.current = distance < 80;
};
el.addEventListener("scroll", onScroll, { passive: true });
return () => el.removeEventListener("scroll", onScroll);
}, []);
useEffect(() => {
const el = ref.current;
if (el && pinned.current) {
el.scrollTop = el.scrollHeight;
}
}, [dependency]);
return ref;
}
The scroll listener records whether the user is near the bottom. After each update, the second effect scrolls only if they were. Using a ref instead of state avoids re-rendering on every scroll event.
The Chat Component
Now the pieces come together. The input is a form, so Enter submits naturally, and Shift+Enter inserts a newline:
// src/chat/Chat.tsx
import { useState, type FormEvent, type KeyboardEvent } from "react";
import { useChat } from "./useChat";
import { useStickToBottom } from "./useStickToBottom";
import { MessageBubble } from "./MessageBubble";
export function Chat() {
const { messages, isStreaming, send, stop } = useChat();
const [draft, setDraft] = useState("");
const scrollRef = useStickToBottom<HTMLDivElement>(messages);
const submit = (event?: FormEvent) => {
event?.preventDefault();
if (!draft.trim() || isStreaming) return;
void send(draft);
setDraft("");
};
const onKeyDown = (event: KeyboardEvent<HTMLTextAreaElement>) => {
if (event.key === "Enter" && !event.shiftKey && !event.nativeEvent.isComposing) {
event.preventDefault();
submit();
}
};
const lastMessage = messages.at(-1);
return (
<div className="mx-auto flex h-dvh max-w-3xl flex-col">
<div ref={scrollRef} className="flex-1 space-y-4 overflow-y-auto p-4" aria-busy={isStreaming}>
{messages.length === 0 && (
<p className="mt-20 text-center text-slate-500">Ask anything about React.</p>
)}
{messages.map((m) => (
<MessageBubble key={m.id} message={m} />
))}
</div>
<p className="sr-only" aria-live="polite">
{!isStreaming && lastMessage?.role === "assistant" && lastMessage.status === "done"
? "Reply received."
: ""}
</p>
<form onSubmit={submit} className="flex gap-2 border-t p-4">
<label htmlFor="chat-input" className="sr-only">
Message
</label>
<textarea
id="chat-input"
value={draft}
onChange={(e) => setDraft(e.target.value)}
onKeyDown={onKeyDown}
rows={2}
placeholder="Type a message. Shift+Enter for a new line."
className="flex-1 resize-none rounded-lg border p-2"
/>
{isStreaming ? (
<button type="button" onClick={stop} className="rounded-lg border px-4">
Stop
</button>
) : (
<button type="submit" disabled={!draft.trim()} className="rounded-lg bg-blue-600 px-4 text-white disabled:opacity-50">
Send
</button>
)}
</form>
</div>
);
}
Two accessibility choices are worth explaining. The isComposing check stops Enter from sending while someone is typing with an input method editor, which is common for Japanese, Chinese, and Korean users. And instead of putting the streaming text in a live region, which would make screen readers read out every fragment, a separate visually hidden region announces once when the reply is complete.
Handling Fast Streams Efficiently
Each text event triggers a state update. React batches updates that happen in the same task, but stream chunks arrive in separate tasks, so a fast model can cause dozens of renders per second, each re-parsing the growing Markdown. For most chats that's fine. If profiling shows it isn't, buffer chunks and flush them once per animation frame:
// Inside useChat, replacing the direct update for text events
let pending = "";
let frame = 0;
const flush = () => {
frame = 0;
const text = pending;
pending = "";
updateMessage(assistantId, (m) => ({ ...m, content: m.content + text }));
};
// In the event loop:
if (event.type === "text") {
pending += event.text;
if (!frame) frame = requestAnimationFrame(flush);
}
Call cancelAnimationFrame(frame) and flush() once more when the stream ends so the last chunk isn't lost. This caps rendering at the display refresh rate no matter how fast tokens arrive.
Common Mistakes When Streaming Chat Responses
- Calling the model API from the browser. It exposes your API key. Always proxy through your own server.
- Using a non-functional state update in the stream loop. Each chunk overwrites the last with stale state and text disappears.
- Splitting on chunks instead of lines. Network chunks don't respect your message boundaries. Buffer until a full line arrives.
- Not aborting on the server. If the browser disconnects but the server keeps streaming from the model, you pay for output nobody sees.
- Proxy buffering. Responses arrive all at once in production but stream locally. Disable buffering in your reverse proxy or hosting platform for this route.
- Rendering model output as raw HTML. Use a Markdown renderer that builds elements, and never pass model output to
dangerouslySetInnerHTMLwithout sanitizing. - Auto-scrolling unconditionally. Users who scroll up to reread lose their place on every chunk.
Frequently Asked Questions (FAQ) About Streaming Chat UIs in React
Either works. The browser's EventSource API only supports GET requests, which makes sending a conversation awkward, so many apps use fetch with a streamed response body instead. You can still send data in SSE format over fetch and parse it yourself. NDJSON is simpler if you control both ends.
Something between the server and the browser is buffering. Common causes are a reverse proxy like Nginx, compression middleware that waits for a full chunk, or a hosting platform that doesn't support streaming for that route type. Disable buffering and compression for the endpoint and check your platform's streaming docs.
Create an AbortController for each request and pass its signal to fetch. Calling abort() cancels the request in the browser. On the server, listen for the response's close event and abort the model stream too, so generation actually stops.
Libraries like the AI SDK provide a ready-made useChat hook, stream protocols, and provider adapters, which saves time on larger projects. Building it yourself, as in this post, is worth doing once so you understand what those abstractions handle, and it's a fine choice for simple apps with one provider.
react-markdown lets you override the code component with the components prop. Pass code blocks to a highlighter such as Shiki or Prism there. Since highlighting the growing text on every chunk is expensive, consider highlighting only after the message status becomes done.
Conclusion
A streaming chat UI comes down to a few well-defined pieces: a server endpoint that keeps your key private, validates input, and forwards model output as a simple event stream; a client reader that decodes bytes and buffers partial lines; a hook that appends chunks with functional state updates; and components that render Markdown, follow the conversation without hijacking scroll, and let users stop generation.
From here, add persistence for conversations, a retry button on errored messages, and code highlighting once a reply finishes. If the app grows, look at moving the chat state into a store or adding tool use on the server, but keep the same structure: the server owns the model, the client owns the stream and the UI.


