HTTP Interface Connection Pooling: Transitioning to Distributed Load Balancing
24.5K reputation · 05 Aug 2023, 21:16 UTC
ClickHouse manages server-side connection limits through the max_connections setting, while the HTTP interface relies on the Connection: keep-alive header and keep_alive_timeout to maintain persistent sessions. In single-node environments, this mechanism handles basic connection reuse effectively.
As deployments scale to distributed clusters, relying on internal server-side pooling becomes problematic. High traffic spikes can lead to connection exhaustion or memory overhead due to the thread-per-connection model, necessitating a transition toward external load balancers like Nginx or HAProxy to manage traffic distribution and socket persistence.
There is uncertainty regarding the optimal alignment of keep_alive_timeout between the ClickHouse server and an external load balancer to prevent premature socket closure or "broken pipe" errors.
- What is the recommended configuration for
keep_alive_timeoutwhen using an external load balancer to ensure connection stability? - How does the server-side
max_connectionslimit interact with the concurrent connection pools maintained by a distributed load balancer?