Reducing Microservice Latency with OpenTelemetry and Jaeger Tracing
Learn how to implement distributed tracing in Java microservices using OpenTelemetry and Jaeger to identify latency bottlenecks and visualize request flows.
07 Jul 2025, 00:16 UTC

The Problem: The \"\\\"Black Box\\\" of Distributed Requests
\nIn a microservices architecture, a single user request often triggers a chain of calls across multiple services. When a request becomes slow, standard application logs only show that a specific service took too long; they do not show a why or where the bottleneck occurred in the wider chain. This lack of visibility makes it impossible to determine if the delay is caused by a slow database query in Service B, a network timeout in Service C, or an inefficient loop in Service A.
\n\nThe solution is distributed tracing. By implementing the OpenTelemetry Java SDK and sending data to a Jaeger backend, you can visualize the entire lifecycle of a request as a series of spans (individual units of work with start and end times) linked by a single Trace ID.
\n\nPrerequisites
\n- \n
- Java 11 or higher. \n
- A running Jaeger instance. For local testing, use the
jaegertracing/all-in-oneDocker image. \n - A Maven or Gradle project with microservices communicating via HTTP. \n
Implementation Strategy: Auto-Instrumentation vs. Manual SDK
\nYou have two primary paths for integrating tracing. The Java Agent provides automatic instrumentation for common libraries (Spring Boot, JDBC, OkHttp) without code changes. The OpenTelemetry SDK allows for manual instrumentation, which is necessary for tracking specific business logic or custom internal methods.
\n\nOption 1: Automatic Instrumentation (Fastest Setup)
\nRun your application with the OpenTelemetry Java agent. This intercepts calls to supported libraries and automatically handles context propagation—the process of passing the Trace ID through HTTP headers (W3C Trace Context) to the next service.
\n\n# Run on the application server\njava -javaagent:opentelemetry-javaagent.jar \\
-Dotel.service.name=order-service \\
-Dotel.exporter.otlp.endpoint=http://jaeger-collector:4317 \\
-jar order-service.jar\n\nOption 2: Manual Instrumentation for Business Logic
\nWhen you need to track a specific complex calculation or a proprietary algorithm, use the SDK to create custom spans. Add the opentelemetry-api dependency to your project.
// Example: Tracking a specific business process\nTracer tracer = GlobalOpenTelemetry.getTracer(\"com.example.orders\");\nSpan span = tracer.spanBuilder(\"calculate-discount\").startSpan();\n\ntry (Scope scope = span.makeCurrent()) {\n // Add custom metadata to the span for easier searching in Jaeger\n span.setAttribute(\"customer.tier\", \"gold\");\n span.setAttribute(\"order.value\", 150.00);\n \n // Your business logic here\n performDiscountCalculation();\n} catch (Exception e) {\n span.recordException(e);\n throw e;\n} finally {\n span.end();\n}\n\nConfiguring Sampling to Prevent Performance Degradation
\nTracing every single request in a high-traffic production environment can saturate your network and overwhelm the Jaeger storage backend. You must configure a sampling strategy to determine which traces are recorded.
\n\n| Strategy | \nConfiguration | \nUse Case | \n
|---|---|---|
| Constant | \nOTEL_TRACES_SAMPLER=always_on | \n Development and QA environments. | \n
| Probabilistic | \nOTEL_TRACES_SAMPLER=parentbased_always_off (with ratio) | \n Production; captures a percentage (e.g., 1%) of traffic. | \n
| Ratio | \nOTEL_TRACES_SAMPLER_ARG=0.1 | \n Capturing 10% of requests to balance detail and overhead. | \n
Verification and Diagnostics
\nTo verify the implementation, follow these steps:
\n- \n
- Trigger a Request: Send a request that traverses at least two services (e.g., Gateway → Order Service → Inventory Service). \n
- Locate the Trace: Open the Jaeger UI (usually
http://localhost:16686), select the service name from the dropdown, and click \"Find Traces\". \n - Check for Fragmentation: If you see multiple separate traces for one request instead of one continuous timeline, your context propagation is failing. Ensure both services are using the same propagation format (W3C is the default for OpenTelemetry). \n
- Inspect Metadata: Click on a custom span to verify that the
customer.tieror other attributes you added manually are visible in the tags section. \n
Limitations and Risks
\n- \n
- Memory Overhead: The Java Agent adds a small amount of memory overhead. Monitor JVM heap usage after deployment. \n
- Network Latency: While the SDK exports spans asynchronously, a massive volume of spans can still impact network throughput. Always use probabilistic sampling in production. \n
- Storage Growth: Jaeger's backend storage (Elasticsearch or Cassandra) can grow rapidly. Implement a TTL (Time-to-Live) policy for trace data. \n
Rollback Procedure
\nBecause this implementation primarily involves adding a Java agent or a library dependency, rollback is straightforward:
\n- \n
- Remove the
-javaagentflag from the JVM startup arguments. \n - Restart the application service. \n
- If using the SDK, remove the OpenTelemetry dependencies and the manual span code via a git revert. \n
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.