Getting Started with New Relic Distributed Tracing for Java Microservices
Learn how to enable New Relic Distributed Tracing in a Java microservice stack, configure sampling, verify header propagation, and balance observability with cost. Practical steps, code snippets, and troubleshooting tips included.
18 Dec 2025, 04:36 UTC

Why Distributed Tracing Matters in a Java Microservice Stack
In a typical microservice architecture, a single user request can touch dozens of services. When a latency spike occurs, you need to know *where* the delay happened: is it a slow database call, a network hiccup, or a mis‑configured service? Traditional APM metrics give you per‑service averages, but they don’t show the request’s journey.
New Relic Distributed Tracing stitches together the individual spans that make up a request flow, giving you a visual “trace graph” that reveals the exact sequence of calls and their latencies.
Prerequisites Before You Flip the Switch
- The New Relic Java agent (v8.0+). Add it to your JVM arguments:
-javaagent:/path/to/newrelic.jar. - A New Relic account on a tier that supports Distributed Tracing (Pro or higher). Check the APM section of your billing dashboard.
- Your services must expose HTTP endpoints (or gRPC) and use a client library that allows header injection.
1. Enable Distributed Tracing in the Agent
Open newrelic.yml in the root of each service and add:
distributed_tracing:
enabled: true
# Optional: set a lower sampling rate to control cost
# sampling_rate: 0.02 # 2%
Restart the application after saving the file. The agent will now automatically inject the traceparent header into outgoing HTTP requests.
2. Propagate the Context Manually (When Needed)
Most HTTP client libraries (e.g., OkHttp, Apache HttpClient) let you add headers programmatically. If you’re using a custom client that doesn’t automatically forward the agent’s context, copy the header from the incoming request:
// Inside a Spring @RestController method
@RequestHeader("traceparent") String traceparent,
...
// When calling downstream service
HttpRequest request = HttpRequest.newBuilder()
.uri(URI.create("http://service-b/endpoint"))
.header("traceparent", traceparent) // propagate
.build();
Verify that the header name matches exactly – the agent expects traceparent (lowercase). A mismatch breaks continuity and results in orphaned spans.
3. Tune the Sampling Rate to Balance Insight and Cost
New Relic stores every trace in its backend. A 100 % sampling rate can quickly inflate ingestion costs, especially in high‑traffic services. The agent’s default is 1 %. If you need more detail, increase it gradually while monitoring the Ingestion tab in the New Relic UI.
# 5 % sampling
distributed_tracing:
enabled: true
sampling_rate: 0.05
Remember: the sampling rate is applied *per service*. If you have 10 services each at 5 %, you’ll still ingest roughly 50 % of the total traffic.
4. Verify Trace Continuity
- UI Check: Open the APM dashboard, click a transaction, and look for a “Trace” tab. A linked graph should display spans from Service A → Service B → Service C.
- Header Inspection: Use a debugging proxy (e.g., Wireshark, mitmproxy) to verify that outgoing requests contain
traceparentand that the value matches the incoming header. - NRQL Validation:
This query lists transactions that have trace IDs, confirming that the agent recorded them.SELECT count(*) FROM Transaction WHERE appName = 'Service-A' SINCE 1 hour ago FACET traceId WHERE traceId IS NOT NULL LIMIT 10
5. Common Pitfalls and How to Avoid Them
- Header Name Mismatch: The agent uses
traceparent. Some libraries default toTraceparentorTRACESTATE. Ensure case‑sensitivity matches. - Excessive Sampling: Setting 100 % on a busy service can double your ingestion cost. Start low, monitor, and scale.
- Non‑instrumented Endpoints: Exclude health checks or metrics endpoints from tracing to reduce noise. In
newrelic.ymlyou can add:transaction_tracer: exclude: - /health - /metrics - Tier Compatibility: Distributed Tracing is not available on the free tier. Verify your plan before enabling.
6. Trade‑Offs: Visibility vs. Cost
The primary trade‑off is between data granularity and ingestion cost. A 1 % sample gives a statistically sound view of typical request paths with minimal cost. If you need to debug a rare edge case, temporarily bump the sampling rate on the affected service, run a controlled load test, then revert.
Another consideration is performance overhead. Benchmarks show the agent adds ~1 ms per request when tracing is enabled. For latency‑sensitive services, measure the impact in a staging environment before rolling out to production.
Actionable Next Steps
- Deploy the agent with
distributed_tracing.enabled: trueon all services. - Verify header propagation in a staging environment.
- Set the sampling rate to 1 % and monitor ingestion.
- Use the New Relic UI to explore trace graphs and identify bottlenecks.
- For critical services, incrementally increase sampling while watching the billing dashboard.
- Document the tracing strategy in your architecture guide so new developers can follow the pattern.
By following these steps you’ll have end‑to‑end visibility into your Java microservices, enabling faster root‑cause analysis and more informed performance tuning—all while keeping costs predictable.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.