The stream dies after 60 seconds

A stream that dies on the same second is a timeout, not the network. Default values in nginx, ALB, Cloudflare, Heroku and Envoy, and why a heartbeat is safer.

What it means

The stream works, events arrive, then the connection closes, always on the same second, usually the sixtieth. Networks break at random. Timeouts fire on a schedule, so this is a timeout; the only question is whose.

How to fix it

Raise the timeout on every proxy you control. Defaults below.

nginx sets proxy_read_timeout to 60 seconds. It counts the gap between two reads from your backend, not the age of the connection. So a stream with an event every ten seconds lives forever, and a stream that goes quiet for a minute dies.

location /events {
    proxy_pass http://app;
    proxy_http_version 1.1;
    proxy_buffering off;
    proxy_read_timeout 3600s;
    proxy_send_timeout 3600s;
    send_timeout 3600s;
}

Three directives, because they count three different things. proxy_read_timeout is a pause in the backend's response. proxy_send_timeout is a pause while writing to the backend. send_timeout is a pause while writing to the client. All three default to 60 seconds.

Kubernetes nginx ingress takes the same numbers as annotations:

nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"

AWS ALB holds its idle_timeout.timeout_seconds attribute at 60 seconds.

aws elbv2 modify-load-balancer-attributes \
  --load-balancer-arn <arn> \
  --attributes Key=idle_timeout.timeout_seconds,Value=3600

Apache inherits ProxyTimeout from Timeout, which is 60 seconds in 2.4. Set ProxyTimeout 3600.

Envoy deserves a second read. The route timeout is 15 seconds by default and it caps the whole response, not the pauses in it. A heartbeat cannot beat it. Set it to 0s and lean on stream_idle_timeout instead, which is 5 minutes.

Cloudflare returns 524 if your origin goes quiet for 125 seconds. Only Enterprise can raise that (up to 6000 s). Everyone else needs a heartbeat.

Heroku is not configurable. Your app has 30 seconds to send something, then 55 seconds between every byte after that, or the router cuts the connection (H12, H15, H28).

gunicorn sets timeout to 30 seconds, and it is not a proxy timeout: the master kills a worker that has produced nothing for that long, and with sync workers a long stream looks exactly like that. Use gevent or eventlet workers, or --timeout 0.

uvicorn has no request timeout at all. --timeout-keep-alive (5 seconds) only covers idle keep-alive connections, so there is nothing here to blame for a cut stream.

A heartbeat beats the setting

You can only raise the timeout on proxies you know about. Between your nginx and the user's browser there may be a corporate proxy, a home router and a mobile carrier. Each has a timeout of its own, and you will never see their config. RFC 9112 § 9.5 lets any of them drop an idle connection at any time.

So the rule is: never go quiet for long. SSE has a comment line for this. It starts with a colon, produces no event, and the client ignores it.

const heartbeat = setInterval(() => res.write(': ping\n\n'), 15_000);
req.on('close', () => clearInterval(heartbeat));

Every 15 seconds, not every 59: timeouts in the proxy chain are sometimes shorter than a minute. The cost is eight bytes.

One catch. A heartbeat only beats idle timeouts. A limit on total duration ends the response no matter how much data you send: that is what Envoy's route timeout does, and plenty of API gateways behave the same way. Those you have to raise by hand.

While you are there, tell the client how long to wait before reconnecting. EventSource applies it on its own:

retry: 3000

Telling a timeout apart from everything else

What you see What it is
Cut on the same second every time a timeout on one of the proxies
Cut at a random time, every client at once the application restarted or crashed
Cut at a random time, one client that client's network
Data arrives in a batch, then a cut buffering plus a timeout

WebSocket does the same thing for the same reason. It just surfaces as close code 1006, which tells you nothing.

See it on your own traffic

You need two numbers: when the connection opened and when it closed. A proxy logs the duration of every request: a column full of exactly sixty seconds means a timeout. The same view shows whether data was still flowing before the cut: if the last event arrived at second fifty-nine, this is not a timeout and the application is where to look. See Capture & Inspect Traffic for how Solpuga records a stream.

View in Solpuga

Tool for this page

SSE stream parserPaste a raw Server-Sent Events stream and see the events, ids, retry hints and framing mistakes.

Related