Replacement characters «�» instead of text in a stream
A multi-byte character was cut at a chunk boundary. One decoder and one flag fix it, and the same mistake is hiding one level up, in the SSE parsing.
What it means
Events arrive, but some characters come out as �, which is U+FFFD, the replacement
character. Short responses are clean, long ones get one or two per paragraph.
The server's encoding is fine. The chunk boundary is not.
How to fix it
In UTF-8 every non-ASCII character takes 2 to 4 bytes
(RFC 3629 § 3 allows 1 to 4
octets per character). A chunk ends where TCP
ended it, not where a character ended. Cut «я» in half and the first chunk ends
with 0xD1, the next starts with 0x8F. Decode each chunk on its own and you
get two �: one for the unfinished sequence, one for the continuation with no
start. A split 4-byte emoji can give three. ASCII text never shows this, which
is why the bug reaches production.
The fix is one decoder for the whole stream plus stream: true:
const decoder = new TextDecoder(); // ONE for the whole stream
const reader = response.body.getReader();
while (true) {
const { value, done } = await reader.read();
if (done) break;
handle(decoder.decode(value, { stream: true })); // ← the flag
}
const tail = decoder.decode(); // "" normally
if (tail) console.warn('stream ended mid-character');
stream: true makes the decoder hold back unfinished bytes until the next
call. new TextDecoder() inside the loop defeats the point, because the state lives in
the instance, and a fresh one remembers nothing.
That last decode() with no arguments is not a flush of pending text. A
finished character is handed over immediately, so there is never any text
waiting. It returns "" for a clean stream and "�" for one that was cut
mid-character. It checks for truncation, nothing more.
TextDecoderStream is shorter and harder to get wrong, and it does the same:
const reader = response.body
.pipeThrough(new TextDecoderStream())
.getReader();
while (true) {
const { value, done } = await reader.read();
if (done) break;
handle(value);
}
for await (const chunk of stream) reads better, but async iteration over a
ReadableStream is new: Chrome 124, Firefox 110, Safari 27. TextDecoderStream
itself goes back to Safari 14.1, so the reader loop is the portable form.
Node runs the same code, because TextDecoder is global there. The Node spelling of
this bug is chunk.toString() on every chunk; StringDecoder from
node:string_decoder is the older fix and keeps the same state.
With EventSource the problem cannot happen: the browser decodes for you. It
hits people reading SSE or NDJSON through fetch, which is everybody who needs
headers or a POST body, neither of which EventSource supports.
The other half of the same bug
Chunk boundaries do not line up with event boundaries either.
data: {"a":1}\n\n can arrive in two pieces, and «split the chunk on \n\n»
loses half of it. The same trick applies: a buffer that outlives the chunk.
let buffer = '';
function handle(text) {
buffer += text;
if (buffer.endsWith('\r')) return; // \r or \r\n? wait for the next chunk
buffer = buffer.replace(/\r\n|\r/g, '\n'); // SSE allows \n, \r\n and bare \r
const parts = buffer.split('\n\n');
buffer = parts.pop() ?? ''; // the tail waits for next time
for (const part of parts) emit(part);
}
The parts.pop() is not optional. The last element is either an empty string
(the chunk ended exactly on a separator) or an unfinished event. Both belong in
the buffer, not in the handler. At done, emit whatever is still in it.
Normalise line endings on the buffer, not on the chunk. \r and \n can
land in different chunks, so a per-chunk text.replace(/\r\n/g, '\n') leaves a
stray \r behind. That is the same boundary bug, one layer up.
The BOM (RFC 3629 § 6) needs
less work than people think. TextDecoder and
TextDecoderStream strip a leading BOM already (ignoreBOM defaults to
false, which means «skip it»). Strip it by hand only if you passed
ignoreBOM: true, or decoded with Node's toString('utf8') or StringDecoder,
which both keep it. A surviving turns the first field name into something
the parser does not know, so the first data: line is silently dropped.
See it on your own traffic
A proxy shows the response body as bytes, not as a string the browser has already stitched together. You can see where the chunk boundary actually falls and whether it landed inside a character. That settles the question: bad encoding from the server, or a decoder that joined the chunks wrong. See Capture & Inspect Traffic for how Solpuga records a stream.
Tool for this page
SSE stream parserPaste a raw Server-Sent Events stream and see the events, ids, retry hints and framing mistakes.Related
- SSE arrives in one chunk: nginx bufferingServer-Sent Events land all at once instead of streaming. Three layers hold them: proxy_buffering, gzip and the app itself. How to test each.
- The stream dies after 60 seconds: proxy_read_timeoutA stream that dies on the same second is a timeout, not the network. Default values in nginx, ALB, Cloudflare, Heroku and Envoy, and why a heartbeat is safer.