Agent fails to initialize after upgrade - Java APM agent
0 reputation · 19 May 2026, 07:15 UTC
New Relic APM agents have a documented post-upgrade failure mode where the agent fails to initialize or fails to report while the instrumented application continues to run. Errors are captured in the agent log rather than crashing the host process. This applies to per-language agents such as the Java agent distributed as newrelic.jar with newrelic.yml, Node.js agent as an npm package, and the Infrastructure agent with its OS package and configuration.
Safe recovery is tied to retaining the prior agent artifact and matching configuration, and to the documented disable switch that can silence the agent without uninstalling it. Config drift between an old configuration file and new agent defaults can change behavior after an upgrade. Rollback restores a known-good telemetry state but leaves the deployment on an older version subject to end-of-life. Roll forward preserves version currency but requires diagnosis with degraded or absent telemetry.
The unresolved decision is which path to commit to for a live service after a failed upgrade.
What decision criteria distinguish a rollback to the previous artifact from a roll-forward with the upgraded agent when telemetry is degraded? Which configuration compatibility checks are required before choosing either path? When is disabling the agent appropriate in the decision without extending the telemetry gap?