The stream dies after 60 seconds
A stream that dies on the same second is a timeout, not the network. Default values in nginx, ALB, Cloudflare, Heroku and Envoy, and why a heartbeat is safer.
What it means
The stream works, events arrive, then the connection closes, always on the same second, usually the sixtieth. Networks break at random. Timeouts fire on a schedule, so this is a timeout; the only question is whose.
How to fix it
Raise the timeout on every proxy you control. Defaults below.
nginx sets proxy_read_timeout to 60 seconds. It counts the gap between two
reads from your backend, not the age of the connection. So a stream with an event
every ten seconds lives forever, and a stream that goes quiet for a minute dies.
location /events {
proxy_pass http://app;
proxy_http_version 1.1;
proxy_buffering off;
proxy_read_timeout 3600s;
proxy_send_timeout 3600s;
send_timeout 3600s;
}
Three directives, because they count three different things. proxy_read_timeout
is a pause in the backend's response. proxy_send_timeout is a pause while
writing to the backend. send_timeout is a pause while writing to the client.
All three default to 60 seconds.
Kubernetes nginx ingress takes the same numbers as annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
AWS ALB holds its idle_timeout.timeout_seconds attribute at 60 seconds.
aws elbv2 modify-load-balancer-attributes \
--load-balancer-arn <arn> \
--attributes Key=idle_timeout.timeout_seconds,Value=3600
Apache inherits ProxyTimeout from Timeout, which is 60 seconds in 2.4. Set
ProxyTimeout 3600.
Envoy deserves a second read. The route timeout is 15 seconds by default and
it caps the whole response, not the pauses in it. A heartbeat cannot beat it.
Set it to 0s and lean on stream_idle_timeout instead, which is 5 minutes.
Cloudflare returns 524 if your origin goes quiet for 125 seconds. Only Enterprise can raise that (up to 6000 s). Everyone else needs a heartbeat.
Heroku is not configurable. Your app has 30 seconds to send something, then 55 seconds between every byte after that, or the router cuts the connection (H12, H15, H28).
gunicorn sets timeout to 30 seconds, and it is not a proxy timeout: the master
kills a worker that has produced nothing for that long, and with sync workers a
long stream looks exactly like that. Use gevent or eventlet workers, or
--timeout 0.
uvicorn has no request timeout at all. --timeout-keep-alive (5 seconds) only
covers idle keep-alive connections, so there is nothing here to blame for a cut
stream.
A heartbeat beats the setting
You can only raise the timeout on proxies you know about. Between your nginx and the user's browser there may be a corporate proxy, a home router and a mobile carrier. Each has a timeout of its own, and you will never see their config. RFC 9112 § 9.5 lets any of them drop an idle connection at any time.
So the rule is: never go quiet for long. SSE has a comment line for this. It starts with a colon, produces no event, and the client ignores it.
const heartbeat = setInterval(() => res.write(': ping\n\n'), 15_000);
req.on('close', () => clearInterval(heartbeat));
Every 15 seconds, not every 59: timeouts in the proxy chain are sometimes shorter than a minute. The cost is eight bytes.
One catch. A heartbeat only beats idle timeouts. A limit on total duration
ends the response no matter how much data you send: that is what Envoy's route
timeout does, and plenty of API gateways behave the same way. Those you have to
raise by hand.
While you are there, tell the client how long to wait before reconnecting.
EventSource applies it on its own:
retry: 3000
Telling a timeout apart from everything else
| What you see | What it is |
|---|---|
| Cut on the same second every time | a timeout on one of the proxies |
| Cut at a random time, every client at once | the application restarted or crashed |
| Cut at a random time, one client | that client's network |
| Data arrives in a batch, then a cut | buffering plus a timeout |
WebSocket does the same thing for the same reason. It just surfaces as close code 1006, which tells you nothing.
See it on your own traffic
You need two numbers: when the connection opened and when it closed. A proxy logs the duration of every request: a column full of exactly sixty seconds means a timeout. The same view shows whether data was still flowing before the cut: if the last event arrived at second fifty-nine, this is not a timeout and the application is where to look. See Capture & Inspect Traffic for how Solpuga records a stream.
Tool for this page
SSE stream parserPaste a raw Server-Sent Events stream and see the events, ids, retry hints and framing mistakes.Related
- SSE arrives in one chunk: nginx bufferingServer-Sent Events land all at once instead of streaming. Three layers hold them: proxy_buffering, gzip and the app itself. How to test each.
- WebSocket closed with code 1006 — how to find the causeThe connection died with no Close frame. Four causes, and how to tell them apart.