Missing Distributed Traces in Datadog APM: Diagnostic Guide for Microservices
Traces appear for some services but distributed traces are broken. Use this diagnostic guide to check sampling, service names, agent health and intake connectivity in Datadog APM for microservices.
29 Mar 2026, 20:06 UTC

Traces appear in Datadog APM for some services but distributed traces are broken. A request starts in an API gateway, you see its spans, but downstream services show no child spans, latency outliers disappear in Trace Explorer, and trace IDs are inconsistent across services. That is the recognizable condition for incomplete distributed tracing in a Datadog Agent + APM setup.
Recognizable condition
Symptoms cluster around trace linking rather than total absence of data.
- Trace Explorer shows a root span for service A with no children from service B, even though B handled the request.
- Trace IDs differ between services for the same request.
- Service list in APM shows services that never connect to each other.
- Agent status shows APM running but intake queue grows or last submission stalls.
A span is a timed unit of work. A trace is a collection of spans linked by trace_id. Distributed tracing relies on propagation headers and consistent service names to link spans across process boundaries.
Cause / diagnostic table
| Observed symptom | Likely cause | Why it breaks linking |
|---|---|---|
| Child spans missing intermittently | Low sampling rate or dynamic sampling drops child spans | Priority sampling can keep root spans and drop children |
| Services never link despite traffic | Service name mismatch or incorrect service mapping | APM groups spans by service name; mismatches prevent join |
| No spans from a service at all | APM agent not running or not configured for runtime | No exporter to Datadog intake |
| Spans buffered then disappear | Network egress blocked to Datadog intake | Agent buffers and eventually drops |
| Spans accepted but incomplete | High cardinality tags cause rejection or throttling | Ingestion limits drop spans |
Ordered checks
1. Confirm agent health and connectivity
Run on the host where the Datadog Agent is installed. Read-only check.
curl -s http://localhost:5000/statusCheck that the APM check is listed as running and that the intake queue size is stable and last successful submission timestamp is recent. Required permission: local network access to agent. Risk: none for read.
Verify environment variables used by the agent and libraries:
env | grep -E 'DD_AGENT_HOST|DD_API_KEY|DD_APM_ENABLED'Placeholder values should match your Datadog site and agent host. Missing DD_API_KEY or DD_AGENT_HOST prevents intake.
2. Verify sampling configuration
In Datadog organization settings, review APM Sampling. Per-service sampling can be set via library config, e.g.:
DD_TRACE_SAMPLE_RATE=1.0or via Datadog APM sampling rules. Note sampling behavior is version-sensitive across dd-trace libraries and agent releases. Default dynamic sampling rules can change between versions.
Check the APM Services list in Datadog and compare with names emitted by applications. Inconsistent naming is a common root cause.
3. Validate service names and service mapping
Service mapping is used to normalize names across runtimes and to link services when names differ between caller and callee. Check library configuration for service name:
DD_SERVICE=payments-apiCompare the value with the service name shown in APM Services. Service naming conventions differ by language runtime and instrumentation method; custom instrumentation may override automatic naming.
4. Inspect agent logs for intake errors
Run on the agent host with read access to logs.
tail -n 200 /var/log/datadog/apm.logLook for connection warnings to intake endpoints and errors about authentication or payload rejection. Do not share logs containing API keys.
5. Validate trace propagation headers
Confirm that the request headers used for propagation, such as traceparent or datadog, are forwarded between services and not stripped by proxies or API gateways. This is a configuration check in the ingress layer.
Fixes tied to findings
Apply fixes only after the corresponding check confirms the cause.
- Sampling drops children. Increase sampling for critical services using priority sampling to preserve root spans. Example scoped change: set DD_TRACE_SAMPLE_RATE to 1.0 for a critical service during investigation, then revert. Increasing sampling raises ingestion volume and cost; changes should be scoped to critical services and time-boxed.
- Service name mismatch. Align service names across libraries and enable service mapping in Datadog to normalize cross-service naming. Restart the application process after changing DD_SERVICE so the tracer picks up the new name.
- Agent not running or misconfigured. Ensure the agent process is running and APM is enabled. Reconfigure DD_AGENT_HOST and DD_API_KEY to match your site. Restart the agent with appropriate permissions.
- Network egress blocked. Open outbound HTTPS to Datadog intake endpoints for your site. Increase agent buffer settings only as a temporary mitigation. Verify connectivity from the host to intake.
- High cardinality tags. Reduce high cardinality tags and normalize tag keys. Remove request-specific values from tags and use them as span attributes only where needed.
Rollback is relevant for sampling and service name changes. Keep previous values documented so sampling rate and service names can be reverted if cost or naming regressions appear.
Verification
After changes, verify in Datadog APM Trace Explorer: select a trace ID from the originating service and confirm child spans from downstream services are present and linked.
Check the Agent status page again for APM check running, intake queue size and last successful submission timestamp.
Compare service names in APM Services list with names emitted by application logs or library configuration to confirm naming consistency.
Escalation criteria
- Spans are missing across all services despite agent health OK.
- Intake errors persist after network allowlist is verified.
- Sampling changes do not restore trace completeness within one trace retention window.
These indicate potential account limits or platform issue. Escalate with service names, agent version, sampling rules applied, and timestamps of missing traces.
Limitations: This guide assumes Datadog Agent and APM tracing are in use with standard propagation. Custom instrumentation, non-standard proxies, or multi-region setups may require additional checks. Version assumptions apply; verify current behavior in your environment.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.