Guide
Diagnosing High Response Time Alerts in Dynatrace for Microservices
Learn how to triage a Dynatrace high‑response‑time problem by checking CPU, thread pools, downstream latency, GC, and request size, then apply the matching fix.
Published by Tasadduq Burney
15 Jul 2026, 16:53 UTC
3 min146.6K views0

Recognizable condition
Dynatrace opens a problem card when a service’s average response time exceeds the configured threshold (for example, >2 seconds) for five consecutive minutes.
Cause‑symptom table
| Possible cause | Typical symptom | Metric to inspect in Dynatrace |
|---|---|---|
| CPU saturation | High GC pause times, increased request latency | Host CPU usage (%); Process CPU time |
| Thread‑pool exhaustion | Requests queuing, rising response time | Thread pool active / queued counts; Servlet/Jetty/Tomcat thread metrics |
| Downstream dependency latency | External call time grows, overall response time rises | External service call duration; Third‑party call time |
| Garbage‑collection pressure | Frequent long GC pauses, CPU spikes | GC pause time (ms); GC count; Heap usage |
| Large request payloads | Higher processing time, possible network bottlenecks | Request size distribution; Average request bytes |
Ordered diagnostic checks
- Open the service overview page for
and view the CPU and memory usage charts. Look for sustained CPU >80 % or memory pressure. - Navigate to the Thread pools section (or process‑level thread metrics) and check if the active thread count is near the configured maximum and if the queued count is rising.
- In the service flow view, select the outgoing calls to downstream services and examine the average call time and percentile (e.g., p95) for the last 10 minutes.
- Check the GC activity under the process details: GC pause time, GC count, and heap usage trends.
- Open the request attributes or purepath breakdown and review the request size distribution (average size, max size) for the same time window.
Fixes tied to findings
- If CPU saturation is observed – consider horizontal scaling (add instances) or profile hot methods to reduce CPU‑intensive work.
- If thread‑pool exhaustion is seen – increase the pool size in the application configuration or switch to asynchronous processing where appropriate.
- If downstream latency is high – investigate the dependency (check its own Dynatrace metrics), add a circuit‑breaker, or implement request‑level timeout/retry policies.
- If GC pressure is evident – tune the JVM heap size, adjust GC algorithm, or reduce object churn by optimizing allocation patterns.
- If large payloads are the driver – enable response/request compression (e.g., GZIP) or paginate/filter data to reduce transferred bytes.
Escalation criteria
Escalate to the platform or SRE team when:
- Applied remediation does not bring the response time back below the threshold after 15 minutes of observation.
- The problem affects multiple services simultaneously, suggesting a shared infrastructure issue.
- The alert triggers an SLA breach or impacts business‑critical transactions.
Verification steps
- Reproduce the load pattern in a staging environment and confirm that Dynatrace creates a problem matching the condition described.
- After applying a fix, monitor the specific metric you addressed (CPU, thread‑pool usage, downstream latency, GC pause, or request size) to verify it returns to normal levels.
- Ensure the problem card closes automatically and that no related alerts persist for the affected service.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.