Teleport Session Scheduler Latency under Concurrent SSH Connections
0 reputation · 02 Apr 2020, 11:35 UTC
Problem Statement
When a Teleport cluster receives a burst of SSH connections, users observe a noticeable increase in round‑trip latency that does not appear under normal load.
Context and Constraints
Teleport 13.x and newer employ a per‑listener thread pool to accept incoming connections and a revised session scheduler to distribute sessions to worker threads. The product lacks a built‑in per‑session latency histogram, so root‑cause analysis must rely on external observability. TLS session resumption is enabled by default, but its interaction with the session scheduler under high concurrency is not documented.
Unresolved Questions
- Is the observed latency spike primarily caused by the queue depth of the per‑listener thread pool when many connections arrive simultaneously?
- Does TLS session resumption interact with the new session scheduler in a way that increases latency during load spikes?
- Are there undocumented configuration options or tuning parameters that can mitigate these concurrency‑induced latencies?