gRPC keep-alive timeout configuration impact on connection-level timeout behavior
0 reputation · 06 Apr 2022, 02:19 UTC
When operating gRPC inside a service mesh, administrators often configure both per‑channel keep‑alive settings and connection‑level idle timeouts enforced by sidecar proxies or load balancers. The goal is to maintain healthy long‑lived streams while preventing stale connections from consuming resources. However, the interaction between the keep‑alive ping interval and the connection‑level timeout is not clearly defined in the gRPC specification, leading to scenarios where a keep‑alive reset may mask an idle timeout or cause the connection to be considered alive when the mesh has already marked it for termination.
This ambiguity raises questions about which timeout takes precedence, how to tune keep‑alive parameters without unintentionally extending connection lifetimes, and what telemetry is needed to detect ghost connections.
Which timeout—keep‑alive or connection‑level—wins when both are active? How should keep‑alive intervals be adjusted relative to mesh‑enforced idle timeouts to avoid ghost connections? What metrics or logging can reliably indicate that a connection has been kept alive solely by keep‑alive pings?