One Slow Checkout, Three Services: Making Sentry's Distributed Tracing Actually Connect
Sentry traces only help if the header chain survives every hop. A two-service walkthrough, the sampling trade-off, and how to verify propagation.
17 Feb 2026, 18:55 UTC

Three services, one slow request, zero clues
Suppose your checkout flow crosses three services: an API gateway, an orders service, and a payments service. The gateway reports a p95 of roughly 900 ms, but its own work — auth, request parsing, routing — takes 40 ms. The orders service logs a 120 ms database call. The payments service sees nothing slower than 100 ms. Every service looks innocent, and yet the user waits almost a full second.
This is the gap distributed tracing exists to close. Sentry's tracing features stitch those services' work into a single waterfall, so you can see not just what each service did, but what waited on what. Getting that stitched view, though, depends on two things teams often skip: header propagation between services and a deliberate sampling strategy.
How Sentry connects the dots
When an SDK starts a transaction — the root span for one request — every operation inside it (an HTTP call, a database query, a cache read) becomes a span with timing, status, and tags. To make spans from different services share one trace, the calling service sends two headers with its outbound request: sentry-trace, which carries the trace and span IDs plus the sampling decision, and baggage, which carries propagation metadata. Sentry SDKs also emit the standard W3C traceparent header, so traces can interoperate with other observability tooling.
On the receiving side, an SDK that sees an incoming sentry-trace header continues that trace instead of starting a new one. The result is one trace ID spanning every service the request touched. Break the header chain anywhere — an uninstrumented HTTP client, a message queue, a proxy that strips unfamiliar headers — and the trace splits into fragments that look like unrelated requests.
A worked example: gateway to orders
Here is a minimal two-service setup. Run the init code in each service's entrypoint, replace YOUR_DSN with the DSN from your Sentry project settings, and point both services at the same Sentry organization so the spans land in one trace. The API names below reflect the sentry-python 2.x and @sentry/node 8+ generations; they shift between majors, so confirm against your SDK's current docs.
Gateway (Python, sentry-python 2.x):
import sentry_sdk
sentry_sdk.init(
dsn="YOUR_DSN",
traces_sample_rate=0.2, # keep 20% of transactions
)
If the gateway calls the orders service with requests or httpx, the SDK's default integrations inject the trace headers automatically. For an uninstrumented client, inject them yourself:
import requests, sentry_sdk
headers = {
"sentry-trace": sentry_sdk.get_traceparent(),
"baggage": sentry_sdk.get_baggage(),
}
resp = requests.post(
"http://orders.internal/api/orders",
json={"cart_id": cart_id},
headers=headers,
timeout=5,
)
Orders service (Node.js, @sentry/node 8+):
const Sentry = require("@sentry/node");
Sentry.init({ dsn: "YOUR_DSN", tracesSampleRate: 0.2 });
// Initialize before creating the Express app so incoming
// requests are captured and existing traces are continued.
You can also mark business-logic steps with custom spans so they appear in the waterfall:
with sentry_sdk.start_span(op="cache.read", description="cart_lookup"):
cart = cache.get(cart_id)
When propagation works, Sentry's Performance view shows one trace whose waterfall includes the gateway's transaction, the outbound HTTP span, and the orders service's spans nested underneath it, each labeled with its service.
Sampling: one decision, whole-trace consequences
Here is the trade-off that surprises people: the sampling decision is made once, at the root transaction, and propagated downstream. Set traces_sample_rate to 0.2 at the gateway and roughly 20% of full traces are kept — the orders service inherits that decision rather than flipping its own coin. That is good for cost predictability and trace coherence, but it means a rare, pathological request has an 80% chance of never being traced end to end.
If you need finer control, replace the rate with a traces_sampler function that decides per request — for example, always tracing high-value endpoints. Make sure the sampler respects the propagated parent decision where you want continuity, or you will fragment traces across services. Note that error events are reported even when the surrounding transaction is sampled out; you will see the exception, just not the full span timeline around it.
Two more limits before you tag everything: high-cardinality tags (per-user or per-request IDs) on spans inflate storage and degrade search performance, and on self-hosted Sentry, trace volume is a real infrastructure cost — size retention before you raise the sample rate, not after.
Break it on purpose
Before trusting this in production, verify the chain:
- Send one test request through the gateway with curl, then open Sentry's traces view and confirm a single trace contains spans from both services with plausible timings.
- Inspect the call between services — temporarily log the headers or enable the SDK's debug mode — and confirm
sentry-traceandbaggageare present on the outbound request. - Compare trace volume in Sentry against the expected sample rate over a day; a mismatch usually means a hop is dropping headers or a service points at a different DSN.
The five minutes spent proving propagation works is far cheaper than discovering, mid-incident, that your "distributed" trace quietly ended at the gateway.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.