Using Dynatrace Davis AI and SLO Burn Rate Alerts for Faster Microservices Triage
Use Dynatrace Davis AI problem detection and automatic topology with SLO burn rate alerts to triage microservices incidents faster and gate deployments on reliability.
30 Aug 2026, 13:45 UTC

An alert fires at 3am for elevated latency in checkout. The service mesh has 40 services, no one knows which dependency is actually broken, and the on-call engineer is opening traces manually. The useful takeaway is to stop triaging by hunting metrics and instead let automated problem detection build a causal hypothesis and let Service Level Objectives with burn rate alerts act as guardrails for both incident response and deployments.
Problem detection that builds topology for you
Dynatrace uses Davis AI to continuously analyze metrics, logs, traces and events from OneAgent-instrumented hosts and services. OneAgent is Dynatrace’s instrumentation agent that collects process, request flow and telemetry data without manual instrumentation in supported runtimes.
From that data Dynatrace builds an automatic service and dependency topology from process and request flow data. The topology is not a manually maintained diagram; it is derived from observed calls, so impact analysis and blast radius estimation can be shown in the problem view without extra mapping work.
When Davis detects an anomaly it creates a problem and generates causal hypotheses across the full stack. Log and event correlation with custom log parsing can be linked to the problem, so application log anomalies appear in context alongside metrics and traces. This reduces the initial triage step from searching multiple tools to reviewing a single problem with suggested causes.
SLOs and burn rate alerts as engineering guardrails
Service Level Objectives are targets defined on golden signals such as latency, error rate and throughput. Golden signals are the core indicators of service health widely used in SRE practice.
Error budget burn rate alerts measure how fast the error budget is being consumed. The error budget is the amount of failure allowed while still meeting the SLO. A fast burn rate signals a reliability risk that warrants immediate attention, while a slow burn may be acceptable.
Using burn rate alerts as guardrails means deployment pipelines can be configured to pause when the burn rate exceeds a threshold, and on-call triage can prioritize problems that are actively consuming budget versus background noise.
Worked example: payment service SLO with burn rate
Consider a payment service exposed as an HTTP service. An SLO is defined on two golden signals: request latency at the 95th percentile and error rate. The SLO target might be 99.9% of requests succeed with latency under 500ms over a 30 day window.
From that SLO, Dynatrace can derive an error budget and configure burn rate alerts for 1 hour and 6 hour windows. When the burn rate crosses the alert threshold, a problem is created and linked to the service in the topology.
In the problem view the engineer sees the affected service, its immediate dependencies, and Davis causal hypotheses such as increased latency in a downstream fraud-check service correlated with log anomalies. The topology shows blast radius, e.g., checkout and order services are impacted, while catalog is not.
Practical verification steps before relying on this in production include reviewing current Dynatrace product documentation for Davis AI problem detection, causal analysis and service topology capabilities, validating OneAgent supported technologies and OpenPipeline ingestion options against your runtime and cloud environment, and testing SLO definitions and burn rate alert thresholds in a non-production environment before using them for deployment gating.
Limitations and trade-offs
Feature names and UI placement vary by Dynatrace edition and release, and behavior differs between SaaS and Managed deployments. Verify in your environment.
Full-stack instrumentation increases telemetry volume and licensing cost. Sampling and data retention settings affect fidelity and root cause accuracy. Automated root cause suggestions are heuristic and require human validation; noisy environments can produce false positives and alert fatigue.
Do not treat Davis hypotheses as definitive. Use them as a starting point for investigation and confirm with traces and logs.
Actionable closing
Start by ensuring OneAgent coverage for critical services and enabling log ingestion with custom parsing for key applications. Define one SLO per critical service on latency and error rate, add burn rate alerts, and review the problem view during the next incident to assess whether causal hypotheses reduce time to first action. Keep SLO thresholds conservative initially and adjust after observing alert behavior in non-production.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.