Closing Visibility Gaps in Async Workflows with Datadog Custom Spans
Learn how to use Datadog custom spans to bridge visibility gaps in asynchronous workflows and enrich your traces with business-specific metadata.
28 Aug 2026, 16:45 UTC

The 'Black Box' of Asynchronous Processing
Automatic instrumentation is a powerful starting point for observability, but it often fails at the boundary of asynchronous tasks. When a request is handed off to a message queue (like RabbitMQ or Kafka) or a background worker (like Celery or Sidekiq), the trace often breaks. You see the request enter the queue, and you see a separate process pick up a job, but you lose the causal link between the two.
The result is a visibility gap: you know a background job is slow, but you can't tell which specific user request triggered it or what the state of the system was at the moment of dispatch. To fix this, you must move beyond automatic instrumentation and implement custom spans to manually bridge these async boundaries.
Bridging the Trace Gap
A span represents a single unit of work within a trace. While Datadog automatically creates spans for HTTP requests and database queries, custom spans allow you to wrap specific business logic or manual hand-offs. To maintain a single trace across a distributed system, the trace_id and span_id must be propagated through the transport layer (e.g., as metadata in a message queue payload).
Enriching Spans with Business Context
A trace that only shows "ProcessJob" is rarely useful. Custom spans allow you to attach tags—key-value pairs that provide operational context. By adding tags like customer_id or order_type, you can filter your Trace Map to see if latency is isolated to a specific customer segment rather than a global system failure.
Worked Example: Manual Span Implementation (Python)
In this scenario, we are wrapping a heavy data processing function that runs asynchronously. We assume the ddtrace library is installed and the Datadog Agent is running on the host.
from ddtrace import tracer
def process_async_payload(payload):
# Start a custom span to wrap the business logic
with tracer.trace("payload.processing", service="worker-service", resource="process_data") as span:
try:
# Attach business-specific tags for filtering
span.set_tag("customer.id", payload.get("customer_id"))
span.set_tag("payload.size", len(payload.get("data", [])))
# Simulate heavy processing
result = perform_calculation(payload)
span.set_tag("processing.status", "success")
except Exception as e:
span.set_tag("error", True)
span.set_tag("error.msg", str(e))
raise e
Implementation Details
- Execution Environment: Run this within your application code on the worker node.
- Permissions: The application must have network access to the Datadog Agent (typically port 8126).
- Placeholders: Replace
"worker-service"with your actual service name to ensure it groups correctly in the APM UI. - Risk: Wrapping tight loops (e.g., inside a
forloop iterating 10,000 times) with custom spans will create massive overhead and potentially crash the agent or inflate costs.
Trade-offs and Constraints
Custom instrumentation is not a "set and forget" solution. There are two primary risks to manage:
The Cardinality Trap
Avoid using high-cardinality data as tag keys. While the tag value can be unique (like a user_id), creating dynamic keys (e.g., user_123_action: "login") can overwhelm the indexing system and degrade query performance in the Datadog UI.
Over-Instrumentation Overhead
Every span created adds a small amount of CPU and memory overhead. If you instrument every private method in your class, you increase the "observer effect," where the act of monitoring the code slows down the code itself. Focus custom spans on boundaries: API calls, queue hand-offs, and expensive computational blocks.
Verifying the Implementation
To confirm your custom spans are working, follow these steps:
- Trigger the asynchronous process in your environment.
- Navigate to APM > Traces in the Datadog console.
- Search for the specific
servicename used in yourtracer.tracecall. - Click into a trace and verify that the
payload.processingspan appears in the flame graph with the expected custom tags in the side panel.
If spans are missing, check the DD_TRACE_SAMPLE_RATE environment variable on your host; if it is set too low, your custom spans may be sampled out and not sent to the server.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.