Diagnosing Slow Transactions in New Relic APM
A technical guide to diagnosing and fixing slow transactions in New Relic APM, focusing on Apdex drops, database latency, and JVM performance.
08 Sept 2026, 11:47 UTC

Identifying the Performance Degradation
A performance degradation is recognized when your APM dashboard shows an Apdex score dropping below 0.7 or average response times exceeding your defined Service Level Agreement (SLA)—typically over 2 seconds—for a sustained period of 5 minutes or more. Apdex (Application Performance Index) is a standard that measures user satisfaction based on response time thresholds.
Quick Diagnostic Matrix
Use this table to correlate the primary symptom in New Relic with the likely root cause.
| New Relic Symptom | Likely Root Cause | Primary Metric to Check |
|---|---|---|
| High % time in Database | Unoptimized queries or lock contention | Slowest Queries / Explain Plan |
| High % time in External Services | Downstream API latency or timeouts | External Services breakdown |
| High % time in Application Code | Inefficient loops or thread blocking | Transaction Trace / Profiling |
| Spikes in Response Time + CPU | JVM Garbage Collection (GC) or Host saturation | Infrastructure CPU/Memory metrics |
Ordered Diagnostic Workflow
Follow these steps in order to isolate the bottleneck without introducing noise into your analysis.
Step 1: Transaction Breakdown Analysis
Navigate to the Transactions tab in the New Relic UI. Identify the specific transaction with the highest response time and examine the breakdown. Determine if the time is spent in the database, external calls, or the application runtime.
Step 2: Database Query Drill-down
If the database is the bottleneck, go to Query. Identify queries with an average duration exceeding 500ms. Use the Explain plan feature within New Relic to check for full table scans or missing indexes.
Step 3: External Service Validation
If external calls are the cause, review the External Services dashboard. Check for a correlation between increased response times and an increase in error rates (HTTP 5xx) from the downstream provider.
Step 4: Runtime and Host Health
If the breakdown shows high application time but no specific slow query or call, check the Infrastructure integration. Look for:
- JVM Metrics: High GC pause times or escalating thread counts.
- Host Resources: CPU utilization consistently above 80% or memory pressure causing swap usage.
Remediation Strategies
Apply the fix that corresponds to your findings from the workflow above.
Database Bottlenecks
Add missing indexes to columns used in WHERE clauses or refactor complex joins. If lock contention is present, review transaction isolation levels.
External Service Latency
Implement a circuit breaker pattern to prevent cascading failures. Set strict timeouts and retry logic with exponential backoff to avoid overloading the downstream service.
Runtime and Resource Issues
- Thread Contention: Increase the connection pool size or optimize synchronized blocks in the code.
- GC Pauses: Tune the JVM heap size (
-Xmx) or switch to a more performant GC algorithm (e.g., G1GC). Warning: Test these changes in staging first as they can impact overall stability. - Host Saturation: Scale the workload horizontally by adding more pods/instances or scale vertically by increasing CPU/RAM.
Verification and Validation
To verify the fix, run a load test in a lower environment that reproduces the condition. Once deployed to production, use the following NRQL query in the New Relic Query Builder to confirm the average duration has returned to acceptable levels:
SELECT average(duration) FROM Transaction WHERE appName = 'YOUR_APP_NAME' SINCE 10 minutes ago
Ensure that the Apdex score in the UI returns to your target threshold and that associated alert policies have cleared.
Escalation Criteria
Escalate to a P1 incident and notify the on-call SRE if any of the following occur after remediation attempts:
- Apdex remains below 0.7 for more than 15 continuous minutes.
- The transaction error rate exceeds 5%.
- The root cause cannot be isolated via standard APM metrics, requiring the activation of Detailed Distributed Tracing or Profiling.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.