keepalive versus worker_connections: trade‑off for short‑lived connection latency in nginx
0 reputation · 28 May 2021, 03:41 UTC
0 reputation · 28 May 2021, 03:41 UTC
The goal is to lower request latency for short‑lived, bursty traffic in nginx without exceeding system file descriptor limits or causing excessive memory use.
Increasing worker_connections lets each worker handle more simultaneous sockets, but raises memory pressure and may hit FD limits. Enabling keepalive reuses existing TCP connections, cutting handshake overhead, yet its effectiveness depends on keepalive_timeout and upstream timeout settings, and too high a keepalive value can retain idle connections unnecessarily.
Should I prioritize raising worker_connections or configuring keepalive to achieve the lowest latency? What keepalive value balances connection reuse with FD availability given a typical worker_processes auto setting? How does the interaction between keepalive_timeout and upstream server timeouts affect the decision?
26525 reputation · 28 May 2021, 08:23 UTC
For short‑lived, bursty traffic the keepalive_timeout is usually the more effective lever for reducing per‑request latency. A higher timeout keeps the TCP socket open, eliminating the TCP three‑way handshake for subsequent requests from the same client. worker_connections is a safety net: raise it only if you hit the FD ceiling and the keep‑alive pool is already saturated.
keepalive_timeout between 60 s and 120 s if the same client IP makes many rapid requests (e.g., API polling or mobile app traffic). This keeps enough sockets alive to cover the burst without holding them indefinitely.
keepalive_requests to the expected maximum number of requests per burst (default 100 is usually fine). Each keep‑alive socket can serve that many requests before being closed.
worker_connections above the OS file‑descriptor limit (ulimit -n) unless you have verified that memory consumption remains acceptable.
keepalive_timeout is lower than or equal to the upstream server’s timeout (e.g., proxy_read_timeout) so that nginx does not keep a socket open for a connection that the upstream will already have dropped.
In nginx 1.10+ each keep‑alive connection consumes one worker_connections slot. The total number of simultaneous keep‑alive sockets is bounded by worker_connections * (1 + (keepalive_requests‑1)/keepalive_requests). Raising keepalive_timeout increases the number of sockets that stay open, potentially approaching that bound. If the bound is reached, nginx starts closing idle sockets or rejecting new ones, which forces a fresh TCP handshake and increases latency.
ulimit -n
cat /proc/sys/fs/file-max
sudo lsof -p $(pidof nginx) | wc -l
# or use stub_status: http://localhost/nginx_status
worker_connections (e.g., >90% of the limit), consider raising worker_connections after increasing the OS limit.
keepalive_timeout:
http {
keepalive_timeout 90s;
keepalive_requests 100;
# upstream timeouts
proxy_read_timeout 120s;
proxy_connect_timeout 30s;
}
When nginx keeps a socket alive longer than the upstream’s timeout, the upstream may close the connection silently. nginx will then detect a broken socket on the next request and open a new one, adding latency. Therefore set keepalive_timeout to a value that is no greater than the smallest upstream timeout you rely on.
To fine‑tune the recommendation I’d need to know the average number of distinct client IPs per second and the typical burst size per client. If you can share those numbers, I can suggest a more precise keepalive_timeout and whether a worker_connections increase is warranted.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.