# Streaming From Express to React — The Full Picture Most Tutorials Skip

* * *

Here's a scenario that's become pretty standard in 2026: your MERN app calls an LLM, and the user has to sit there staring at a loading spinner for 3–5 seconds before any text appears. The API is streaming the response fine — you're just collecting it all on your Express server and flushing it at the end. That's the problem. Let's fix it, from both ends of the wire.

This post covers the full picture: how to stream from an Express backend and how to correctly consume that stream on the React side. No hand-waving, actual code you can drop in.

* * *

## What "Streaming" Actually Means at the HTTP Level

Before jumping into code, worth being clear on the mechanism. HTTP streaming works through **chunked transfer encoding** — instead of the server sending one big response body, it sends pieces (chunks) over an open connection and signals completion with a zero-length chunk.

There are two common patterns built on top of this:

**Server-Sent Events (SSE)** — A specific text format over `text/event-stream`. The connection stays open, the server pushes named events, and the browser handles reconnection automatically. One-directional: server to client only.

**Raw chunked fetch** — You open a `fetch()` request and read the response body as a `ReadableStream`. More flexible (works with `POST`, lets you send a request body), but you handle the parsing yourself.

SSE is simpler for most streaming use cases. Raw streaming is better when you need POST semantics (like sending a prompt body to your LLM endpoint).

![](https://cdn.hashnode.com/uploads/covers/69d007f5e466e2b7625cd1df/27c0f16c-566b-4332-b7e8-c3755060d9b4.png align="center")

* * *

## Server Side: Streaming SSE From Express

Here's a clean Express endpoint that proxies an LLM stream as SSE:

```javascript
// routes/chat.js
import express from 'express';
import OpenAI from 'openai';

const router = express.Router();
const openai = new OpenAI();

router.get('/stream', async (req, res) => {
  const { prompt } = req.query;

  // SSE headers — these three are required
  res.setHeader('Content-Type', 'text/event-stream');
  res.setHeader('Cache-Control', 'no-cache');
  res.setHeader('Connection', 'keep-alive');

  // Flush headers immediately so the browser knows we're streaming
  res.flushHeaders();

  try {
    const stream = await openai.chat.completions.create({
      model: 'gpt-4o',
      messages: [{ role: 'user', content: prompt }],
      stream: true,
    });

    for await (const chunk of stream) {
      const token = chunk.choices[0]?.delta?.content ?? '';
      if (token) {
        // SSE format: "data: <payload>\n\n"
        res.write(`data: ${JSON.stringify({ token })}\n\n`);
      }
    }

    // Signal the end of the stream
    res.write('data: [DONE]\n\n');
    res.end();
  } catch (err) {
    res.write(`data: ${JSON.stringify({ error: err.message })}\n\n`);
    res.end();
  }

  // Clean up if the client disconnects early
  req.on('close', () => {
    stream?.controller?.abort();
  });
});
```

A few things worth pointing out here. `res.flushHeaders()` is important — without it, some Node.js environments buffer the headers until the first `res.write()`, which can add latency. The `\n\n` double newline in the SSE format is not optional; that's how the browser's `EventSource` parser knows where one event ends and the next begins. And cleaning up on `req.on('close')` is important — if the user navigates away, you want to cancel the upstream LLM request, not burn tokens on a response nobody's reading.

* * *

## Server Side: Streaming a MongoDB Cursor

Same idea applies to large database reads. Instead of loading 50,000 documents into memory, stream the cursor:

```javascript
router.get('/export', async (req, res) => {
  res.setHeader('Content-Type', 'text/event-stream');
  res.setHeader('Cache-Control', 'no-cache');
  res.setHeader('Connection', 'keep-alive');
  res.flushHeaders();

  const cursor = db.collection('orders').find({ status: 'completed' });

  for await (const doc of cursor) {
    res.write(`data: ${JSON.stringify(doc)}\n\n`);
  }

  res.write('data: [DONE]\n\n');
  res.end();
});
```

Your memory footprint stays flat regardless of how many documents exist. This is particularly handy for export features where users want to download or process large datasets.

* * *

## Frontend: Consuming SSE With EventSource

For GET-based SSE endpoints, the browser's built-in `EventSource` is the cleanest option:

```javascript
// hooks/useSSEStream.js
import { useEffect, useRef, useState } from 'react';

export function useSSEStream(url) {
  const [tokens, setTokens] = useState([]);
  const [done, setDone] = useState(false);
  const sourceRef = useRef(null);

  useEffect(() => {
    const source = new EventSource(url);
    sourceRef.current = source;

    source.onmessage = (event) => {
      if (event.data === '[DONE]') {
        setDone(true);
        source.close();
        return;
      }

      const { token } = JSON.parse(event.data);
      setTokens((prev) => [...prev, token]);
    };

    source.onerror = () => {
      source.close();
    };

    return () => source.close();
  }, [url]);

  return { text: tokens.join(''), done };
}
```

Usage:

```jsx
function ChatResponse({ prompt }) {
  const { text, done } = useSSEStream(`/api/stream?prompt=${encodeURIComponent(prompt)}`);

  return (
    <p>
      {text}
      {!done && <span className="cursor-blink">▌</span>}
    </p>
  );
}
```

`EventSource` automatically reconnects if the connection drops, which is nice for flaky networks. The downside: it only works with GET requests. If your endpoint needs a request body (like passing a full conversation history), you'll need the fetch approach below.

* * *

## Frontend: fetch() + ReadableStream for POST Endpoints

This is the pattern you want when you need to POST a body to your streaming endpoint:

```javascript
// hooks/useStreamingFetch.js
import { useCallback, useRef, useState } from 'react';

export function useStreamingFetch() {
  const [output, setOutput] = useState('');
  const [streaming, setStreaming] = useState(false);
  const abortRef = useRef(null);

  const startStream = useCallback(async (messages) => {
    // Cancel any in-flight stream
    abortRef.current?.abort();
    abortRef.current = new AbortController();

    setOutput('');
    setStreaming(true);

    try {
      const res = await fetch('/api/chat/stream', {
        method: 'POST',
        headers: { 'Content-Type': 'application/json' },
        body: JSON.stringify({ messages }),
        signal: abortRef.current.signal,
      });

      const reader = res.body.getReader();
      const decoder = new TextDecoder();

      while (true) {
        const { done, value } = await reader.read();
        if (done) break;

        const chunk = decoder.decode(value, { stream: true });

        // Parse SSE lines manually
        for (const line of chunk.split('\n')) {
          if (!line.startsWith('data: ')) continue;
          const data = line.slice(6).trim();
          if (data === '[DONE]') break;

          try {
            const { token } = JSON.parse(data);
            setOutput((prev) => prev + token);
          } catch {
            // Partial chunk — will reassemble on next read
          }
        }
      }
    } catch (err) {
      if (err.name !== 'AbortError') console.error(err);
    } finally {
      setStreaming(false);
    }
  }, []);

  const stop = useCallback(() => {
    abortRef.current?.abort();
    setStreaming(false);
  }, []);

  return { output, streaming, startStream, stop };
}
```

The `decoder.decode(value, { stream: true })` call is subtle but important — setting `stream: true` tells the `TextDecoder` that this might be a partial character boundary, so it holds incomplete multi-byte characters until the next chunk arrives rather than outputting garbage.

![](https://cdn.hashnode.com/uploads/covers/69d007f5e466e2b7625cd1df/3d13a854-c0af-4ba6-87ff-aa4a5b6c8fd3.png align="center")

* * *

## The AbortController Pattern — Don't Skip This

Notice both examples use `AbortController`. This matters for two reasons: first, it stops the network request when a component unmounts so you don't update state on an unmounted component. Second, it signals your Express server (via `req.on('close')`) to cancel the upstream LLM call.

Without this, every navigation away from your chat page silently keeps a token-burning LLM generation running on the server until it completes naturally. In production that adds up fast.

```jsx
// Clean abort on unmount
useEffect(() => {
  return () => abortRef.current?.abort();
}, []);
```

One line, but it closes a pretty leaky resource loop.

* * *

## Putting It Together in a MERN App

The flow end-to-end looks like this:

1.  React component calls `startStream(messages)`
    
2.  `fetch()` opens a POST to your Express `/api/chat/stream` endpoint
    
3.  Express forwards the request to OpenAI with `stream: true`
    
4.  As tokens arrive from OpenAI, Express writes them as SSE chunks
    
5.  The `ReadableStream` reader on the frontend receives chunks and updates state
    
6.  User sees text appearing token by token — no spinner, no wait
    

If you want a stop button, `stop()` from the hook aborts the fetch, which triggers `req.on('close')` on the server, which aborts the OpenAI stream. The whole chain unwinds cleanly.

* * *

## Wrap Up

Streaming isn't complicated once you see both sides of it. SSE on the server is basically `res.write()` in a loop with the right headers. On the client, `EventSource` handles the simple case and `fetch + ReadableStream` handles anything that needs a POST body. The `AbortController` wires the cancel path together so nothing leaks.

If your app makes any LLM calls or serves large data payloads, this pattern is worth wiring up — the UX difference between "wait 4 seconds, then see everything" and "see text appear immediately" is one of those things users notice immediately, even without being told what changed.
