How to Diagnose AWS Performance Issues Using the Well-Architected Framework’s Measurement Approach
0 reputation · 21 Aug 2020, 07:31 UTC
0 reputation · 21 Aug 2020, 07:31 UTC
How can I apply the AWS Well-Architected Framework’s performance‑efficiency pillar to systematically measure, diagnose, and remediate performance issues in an Amazon EC2‑based application? I am looking for a step‑by‑step explanation that covers baseline establishment, metric selection (such as CPU utilization, network throughput, and latency), data collection tools (like Amazon CloudWatch, Enhanced Monitoring, or X‑Ray), analysis techniques to compare observed values against baseline or best‑practice thresholds, and recommended remediation actions when measurements indicate deviations. Please include any relevant guidance from the Framework or its associated tools that helps ensure the measurement process is repeatable and objective.
Use the Performance Efficiency pillar measurement cycle: define normal, collect at all layers, compare to baseline, remediate, and re-measure. This makes diagnosis repeatable and objective rather than reactive.
Define KPIs tied to business outcomes and capture them under steady-state load. For EC2-based apps typical KPIs are p95/p99 latency, requests per second throughput, error rate %, and resource utilization.
Record thresholds in the AWS Well-Architected Tool as review artifacts so baselines persist across reviews.
The Framework recommends measuring infrastructure, application, and network together.
Assumption: you have CloudWatch enabled on the instances and X-Ray SDK instrumented in the app. If not, first enable collection before analysis.
Compare current dashboards to baseline. Confirmed analysis patterns:
Likely explanation for unexplained latency: external dependencies or micro-bursting not visible in 5-minute averages. Verify with high-resolution metrics and synthetic checks.
Remediation follows measurement.
Validate by re-running the same load profile and confirming KPIs return to baseline. Document changes in the Well-Architected Tool for repeatability.
Use comments to ask for clarification. Post a solution as an answer.
26,525 reputation · 21 Aug 2020, 17:04 UTC
After establishing baseline KPIs, you can create derived metrics with CloudWatch Metric Math to normalize utilization (e.g., CPUUtilization / m1 where m1 is the number of vCPUs) or compute CPU credit utilization for burstable instances (CPUUtilization / CPUCreditBalance). Then enable CloudWatch Anomaly Detection on these metrics to generate dynamic bands that adapt to normal variability. When observed values breach the bands, the deviation is flagged objectively, reducing reliance on static thresholds that may become outdated.
Assumes CloudWatch agent version 1.2.0+ for custom metric publishing and X‑Ray daemon 2.x for tracing.