Diagnosing and Fixing Jenkins Agent Disconnection Issues
A diagnostic guide to resolving Jenkins agent disconnections, covering TCP timeouts, JVM heap exhaustion, and network configuration checks.
08 Mar 2026, 16:04 UTC

The Problem: Intermittent Agent Disconnections
In a distributed Jenkins environment, the "Agent Disconnected" status often appears without a clear failure in the build logs. This usually happens when the Jenkins controller stops receiving heartbeats from the agent, leading the controller to mark the node as offline and aborting all running jobs. The challenge is distinguishing between a network timeout, a resource crash, or a configuration mismatch.
Rapid Diagnostic Table
| Symptom | Likely Cause | Primary Diagnostic Tool |
|---|---|---|
| Immediate disconnect after startup | Version mismatch or Port blockage | nc / telnet |
| Disconnect during heavy build load | JVM Heap exhaustion (GC Pause) | jstat / top |
| Disconnect during idle periods | TCP Timeout / Firewall idle kill | Controller System Log |
| Disconnect during large artifact creation | Disk space exhaustion | df -h |
Step-by-Step Troubleshooting Sequence
1. Verify Network Reachability
Before analyzing logs, confirm the agent can actually reach the controller on the designated agent port (typically a randomly assigned port or a fixed port configured in Global Security). Run this from the agent machine:
# Replace [controller-ip] and [port] with your actual values
nc -zv [controller-ip] [port]
Expected Result: A "Connection succeeded" or "Open" message. If it hangs or returns "Connection refused," check your firewall rules or the Jenkins TCP port for inbound agents.
2. Analyze Agent Process Logs
Check the stdout/stderr of the agent process. Look for specific Java exceptions that indicate why the connection dropped.
java.io.IOException: Connection reset by peer: Usually indicates a firewall or load balancer dropped the connection due to inactivity.java.net.ConnectException: Indicates the agent cannot find the controller.- No logs, but process is dead: This often points to an OS-level OOM (Out of Memory) kill.
3. Check for JVM "Stop-the-World" Pauses
If an agent disconnects only during resource-intensive builds, the JVM may be spending too much time in Garbage Collection (GC), pausing the heartbeat thread long enough for the controller to time it out.
Run jstat on the agent while a build is running to monitor GC behavior:
# Replace [pid] with the agent's Java process ID
jstat -gc [pid] 1000
If the FGC (Full GC) count is rising rapidly and the FGCT (Full GC Time) is increasing significantly, the agent is likely starving for memory.
4. Validate Disk Space and I/O
Jenkins agents write significant amounts of data to the workspace. If the disk fills up, the agent process may hang during I/O operations, leading to a disconnection.
# Run on the agent node
df -h
Ensure the partition hosting the JENKINS_HOME or workspace directory has at least 10-20% free space.
Applying the Fixes
Scenario A: Memory Exhaustion
If jstat confirmed GC pauses, increase the agent's heap size. Modify the agent launch command (or the systemd unit file) to include -Xmx settings.
# Example: Increasing heap to 4GB
java -Xmx4g -jar agent.jar
Risk: If the agent is running in a container (e.g., Kubernetes), ensure the container's memory limit is higher than the JVM heap (e.g., 5GB limit for a 4GB heap) to prevent the OS from killing the process.
Scenario B: TCP Idle Timeouts
If the agent disconnects during idle periods, it is often due to an intermediate firewall killing "silent" TCP connections. Enable TCP Keepalives at the OS level on the agent node:
# Run as root on Linux agent
sysctl -w net.ipv4.tcp_keepalive_time=600
This forces the OS to send a keepalive packet every 10 minutes, preventing the firewall from marking the connection as dead.
Scenario C: Version Mismatch
If the agent fails to connect immediately after a controller upgrade, the agent.jar may be outdated. While Jenkins usually updates this automatically, a corrupted download can cause protocol errors.
- Navigate to
http://[controller]:[port]/jnlpJars/agent.jarin a browser. - Download the jar manually.
- Replace the existing
agent.jaron the agent node and restart the service.
Verification and Rollback
Verification: Go to Manage Jenkins > Nodes. The agent should show as "In sync" with no "Disconnected" warnings in the node log for at least one full build cycle.
Rollback: If increasing JVM heap causes the node to crash (OOMKill), revert the -Xmx value to the previous setting and investigate if the build process itself (e.g., a Maven or Gradle task) is consuming the system memory rather than the Jenkins agent process.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.