Replacement characters «�» instead of text in a stream

A multi-byte character was cut at a chunk boundary. One decoder and one flag fix it, and the same mistake is hiding one level up, in the SSE parsing.

What it means

Events arrive, but some characters come out as �, which is U+FFFD, the replacement character. Short responses are clean, long ones get one or two per paragraph. The server's encoding is fine. The chunk boundary is not.

How to fix it

In UTF-8 every non-ASCII character takes 2 to 4 bytes (RFC 3629 § 3 allows 1 to 4 octets per character). A chunk ends where TCP ended it, not where a character ended. Cut «я» in half and the first chunk ends with 0xD1, the next starts with 0x8F. Decode each chunk on its own and you get two �: one for the unfinished sequence, one for the continuation with no start. A split 4-byte emoji can give three. ASCII text never shows this, which is why the bug reaches production.

The fix is one decoder for the whole stream plus stream: true:

const decoder = new TextDecoder();          // ONE for the whole stream
const reader = response.body.getReader();

while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  handle(decoder.decode(value, { stream: true }));   // ← the flag
}

const tail = decoder.decode();              // "" normally
if (tail) console.warn('stream ended mid-character');

stream: true makes the decoder hold back unfinished bytes until the next call. new TextDecoder() inside the loop defeats the point, because the state lives in the instance, and a fresh one remembers nothing.

That last decode() with no arguments is not a flush of pending text. A finished character is handed over immediately, so there is never any text waiting. It returns "" for a clean stream and "�" for one that was cut mid-character. It checks for truncation, nothing more.

TextDecoderStream is shorter and harder to get wrong, and it does the same:

const reader = response.body
  .pipeThrough(new TextDecoderStream())
  .getReader();

while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  handle(value);
}

for await (const chunk of stream) reads better, but async iteration over a ReadableStream is new: Chrome 124, Firefox 110, Safari 27. TextDecoderStream itself goes back to Safari 14.1, so the reader loop is the portable form.

Node runs the same code, because TextDecoder is global there. The Node spelling of this bug is chunk.toString() on every chunk; StringDecoder from node:string_decoder is the older fix and keeps the same state.

With EventSource the problem cannot happen: the browser decodes for you. It hits people reading SSE or NDJSON through fetch, which is everybody who needs headers or a POST body, neither of which EventSource supports.

The other half of the same bug

Chunk boundaries do not line up with event boundaries either. data: {"a":1}\n\n can arrive in two pieces, and «split the chunk on \n\n» loses half of it. The same trick applies: a buffer that outlives the chunk.

let buffer = '';

function handle(text) {
  buffer += text;
  if (buffer.endsWith('\r')) return;          // \r or \r\n? wait for the next chunk
  buffer = buffer.replace(/\r\n|\r/g, '\n');  // SSE allows \n, \r\n and bare \r
  const parts = buffer.split('\n\n');
  buffer = parts.pop() ?? '';                 // the tail waits for next time
  for (const part of parts) emit(part);
}

The parts.pop() is not optional. The last element is either an empty string (the chunk ended exactly on a separator) or an unfinished event. Both belong in the buffer, not in the handler. At done, emit whatever is still in it.

Normalise line endings on the buffer, not on the chunk. \r and \n can land in different chunks, so a per-chunk text.replace(/\r\n/g, '\n') leaves a stray \r behind. That is the same boundary bug, one layer up.

The BOM (RFC 3629 § 6) needs less work than people think. TextDecoder and TextDecoderStream strip a leading BOM already (ignoreBOM defaults to false, which means «skip it»). Strip it by hand only if you passed ignoreBOM: true, or decoded with Node's toString('utf8') or StringDecoder, which both keep it. A surviving  turns the first field name into something the parser does not know, so the first data: line is silently dropped.

See it on your own traffic

A proxy shows the response body as bytes, not as a string the browser has already stitched together. You can see where the chunk boundary actually falls and whether it landed inside a character. That settles the question: bad encoding from the server, or a decoder that joined the chunks wrong. See Capture & Inspect Traffic for how Solpuga records a stream.

View in Solpuga

Tool for this page

SSE stream parserPaste a raw Server-Sent Events stream and see the events, ids, retry hints and framing mistakes.

Related