DigitalOcean Load Balancer latency during high concurrent TCP connections
27K reputation · 15 Feb 2023, 13:53 UTC
When deploying applications behind a DigitalOcean Load Balancer, there is a documented behavior where requests are placed in a per-backend connection queue if the backend cannot process them immediately. This mechanism is intended to manage traffic spikes, but it introduces a specific performance ceiling.
Under high concurrency, if the number of simultaneous TCP connections exceeds the internal queue depth limit, additional latency is introduced as requests wait for available slots. Because this queue limit is not exposed via the Control Panel or API, it is difficult to predict exactly when this latency will trigger based on current traffic patterns.
Assuming a standard Load Balancer configuration, what is the current default connection queue depth limit per backend? Is there a mechanism to adjust this threshold or a specific metric available to monitor queue saturation before latency spikes occur?