Traefik Mesh timeout propagation and connection pooling behavior
0 reputation · 01 May 2022, 21:58 UTC
0 reputation · 01 May 2022, 21:58 UTC
In a Traefik Mesh environment, request lifecycles are managed via sidecar proxies using Transport resources. While these resources allow for granular control over handshake and response body durations, there is ambiguity regarding how the mesh handles the underlying connection when a mesh-level timeout is triggered.
When a timeout defined in the Traefik Mesh configuration is reached, the sidecar terminates the request. However, it is unclear whether this termination immediately returns the connection to the pool or leaves the connection in a stale state until a socket timeout occurs. This is particularly relevant when application-level socket timeouts differ from mesh-defined settings.
How does Traefik Mesh manage the state of pooled connections after a Transport-level timeout occurs? Is there a mechanism to ensure that the backend service receives a cancellation signal before the sidecar closes the socket?
When a timeout defined in a Traefik Mesh Transport resource is triggered, the sidecar proxy does not return the connection to the pool. Because a timeout at the Transport level indicates a failure in the underlying TCP/TLS handshake or a breach of the response body duration, the connection is marked as invalid and is closed.
Returning a connection to the pool after a timeout would risk "leaking" a corrupted or desynchronized stream into a subsequent request. Consequently, the proxy severs the socket to ensure the next request initiates a fresh handshake.
The mechanism for notifying the backend service of a timeout depends entirely on the protocol being used by the mesh:
RST_STREAM frame. This allows the underlying TCP connection to remain open for other requests while signaling the backend to stop processing that specific stream.There is a functional decoupling between the mesh-defined timeout and the application-level socket timeout. If the Traefik Mesh timeout is shorter than the backend's internal timeout, a "ghost request" scenario occurs: the sidecar returns a timeout error to the client and closes the connection, but the backend service continues to consume CPU and memory resources until its own internal timeout expires or it attempts to send data back to the closed socket.
To confirm how your specific version of Traefik Mesh is handling these terminations, use the following scoped verification methods:
tcpdump on the sidecar container to verify if a FIN or RST packet is dispatched immediately upon the timeout trigger.RUNNABLE state, the cancellation signal is not being propagated effectively.context canceled entries, which indicate the proxy has stopped waiting for the backend.Diagnostic Detail Needed: To provide a more precise recommendation on timeout alignment, please specify if your mesh is utilizing HTTP/1.1 or HTTP/2 for backend communication.
Use comments to ask for clarification. Post a solution as an answer.
No question comments on this page.