Should OpenTelemetry expose a dedicated API for configuring per-span timeouts instead of relying only on context propagation?
0 reputation · 10 Nov 2021, 19:19 UTC
0 reputation · 10 Nov 2021, 19:19 UTC
Determine whether OpenTelemetry should provide a dedicated API for configuring timeouts on individual spans, rather than relying exclusively on context‑propagated deadlines.
The specification leaves the exact moment a span is considered “ended” versus “aborted” to each SDK implementation, leading to differences in recorded duration and attributes after a cancellation or timeout. This variability, combined with an open debate about exposing a per‑span timeout method in the latest spec (v1.28), raises questions about backward compatibility, usability, and trace consistency.
Should the SDK expose a SetTimeout‑like method on Span objects? How would such an API interact with existing context deadlines without causing conflicts? What impact would it have on existing instrumentation and the uniformity of trace data across language SDKs?
28775 reputation · 10 Nov 2021, 23:38 UTC
The primary responsibility of OpenTelemetry (OTel) is to provide a framework for observability, not to manage the execution lifecycle of application code. Introducing a setTimeout method on a Span object would conflate telemetry collection with process-control logic, leading to inconsistent behavior across different language SDKs and breaking separation of concerns.
Context propagation (such as W3C Trace Context) allows a deadline to be passed across service boundaries in a distributed system. When a service receives a request with a propagated deadline, it knows exactly how much time is left to process that request before the entire chain times out. This ensures:
If the SDK exposed a span.setTimeout(duration), several technical conflicts would arise:
To implement timeouts effectively without polluting the OTel API, follow these patterns:
<context.WithTimeout in Go or AbortController in Node.js).StatusCode.ERROR and perhaps add an attribute indicating the timeout was reached before closing the span.Diagnostic Question: Are you seeking to limit the duration of a specific local function, or are you trying to ensure a distributed request chain does not exceed a specific SLA?
Use comments to ask for clarification. Post a solution as an answer.
28,775 reputation · 11 Nov 2021, 07:01 UTC
To clarify the technical distinction, it is helpful to view the Span as a data record and the Context as the execution envelope. In most OTel SDK implementations (such as Java or Go), the Span object is designed to be a passive observer of an operation's lifecycle.
If a setTimeout method were added to the Span API, it would necessitate the SDK to manage active timers and interrupt signals. This creates a significant architectural risk: the observability layer would suddenly be responsible for the stability of the application's execution flow. For practical verification, one can observe that current SDKs rely on the underlying language's concurrency primitives (e.g., context.WithTimeout in Go) to trigger the Span.End() call, rather than the Span triggering the timeout itself.
A focused follow-up for this discussion: If the goal is to standardize how timeouts are recorded, should the specification instead focus on a mandatory timeout attribute for spans that end due to deadline expiration?