Diagnosing and Fixing HPA‑Induced Pod Thrashing in Kubernetes
Diagnose and fix HPA‑induced pod thrashing with a cause‑to‑fix table, step‑by‑step checks, and an example HPA config.
09 Mar 2026, 19:21 UTC

Problem Statement
When a Horizontal Pod Autoscaler (HPA) is misconfigured, it can trigger repeated scale‑up and scale‑down events that leave the cluster in a state of constant churn. This “pod thrashing” wastes resources, increases latency, and can even destabilise other workloads. The most common symptoms are:
- Frequent
SCALE_UPandSCALE_DOWNevents in the HPA event log. - CPU usage hovering near 0 % or 100 % while the HPA keeps adjusting replicas.
- Metrics server reports stale or missing data for the target deployment.
- Pods are repeatedly evicted or recreated.
Below is a diagnostic flow that walks you through identifying the root cause, applying the appropriate fix, and knowing when to involve the cluster‑ops team.
Cause‑to‑Fix Table
| Symptom | Likely Cause | Primary Fix |
|---|---|---|
| HPA scales too often despite low CPU usage | Target CPU utilization set too low | Increase targetCPUUtilizationPercentage or switch to metric‑based scaling |
| Metrics server shows no data for the deployment | Metrics server unhealthy or mis‑configured | Restart metrics‑server deployment or adjust RBAC |
| Pods consume more CPU than requested | Pod resource requests too low | Increase resources.requests.cpu in the deployment spec |
| Cluster autoscaler not adding nodes when needed | Cluster autoscaler disabled or node labels mis‑aligned | Enable autoscaler and correct node pool labels |
| HPA events continue after fixes | Other metrics (memory, custom) drive scaling | Adjust behavior or add metrics spec |
Diagnostic Flow
- Verify HPA configuration
- Run
kubectl get hpa <name> -o yamlin the namespace of the target deployment.kubectl get hpa myapp-hpa -o yaml - Check
spec.targetCPUUtilizationPercentageandspec.minReplicas/maxReplicas. - Confirm the HPA is targeting the correct deployment:
spec.scaleTargetRef.name.
- Run
- Inspect HPA event history
- Run
kubectl describe hpa <name>.kubectl describe hpa myapp-hpa - Look for frequent
SCALE_UPorSCALE_DOWNevents and note the timestamps.
- Run
- Check metrics server health
- Verify the deployment is running:
kubectl get deployment metrics-server -n kube-system.kubectl get deployment metrics-server -n kube-system - Ensure the metrics endpoint is reachable from a pod:
kubectl exec -it <pod> -- curl -s http://metrics-server:4443/metrics. - If the server is down or reports errors, restart it or check its logs.
- Verify the deployment is running:
- Validate pod resource usage
- Use
kubectl top pod <pod> --containersto see real CPU usage versus requests.kubectl top pod myapp-abc123 --containers - Compare the
requests.cpufield in the deployment spec to the observed usage. - If usage consistently exceeds requests, the pod is under‑requested.
- Use
- Examine cluster autoscaler status
- Check if the autoscaler deployment is present and healthy.
kubectl get deployment cluster-autoscaler -n kube-system - Verify node labels match the autoscaler’s
--nodesflag or--scale-down-unneeded-timesettings.
- Check if the autoscaler deployment is present and healthy.
- Review custom metrics (if any)
- If the HPA uses custom metrics, ensure the metrics API is serving data correctly.
- Run
kubectl get --raw /apis/custom.metrics.k8s.io/v1beta1/namespaces/<ns>/podsto list available metrics.
Concrete Example
Below is a minimal HPA that targets CPU utilization. Adjust the targetCPUUtilizationPercentage to match your workload’s typical usage.
apiVersion: autoscaling/v2beta2
kind: HorizontalPodAutoscaler
metadata:
name: myapp-hpa
namespace: default
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: myapp
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
behavior:
scaleUp:
stabilizationWindowSeconds: 30
scaleDown:
stabilizationWindowSeconds: 30
Notice the averageUtilization of 50 %. If your application normally runs at 30 % CPU, the HPA will keep scaling up until the target is met, causing thrashing.
Applying Fixes
- Adjust CPU target
- Increase
averageUtilizationto a realistic value, e.g., 70 %. - Apply the change:
kubectl apply -f myapp-hpa.yaml.
- Increase
- Increase pod resource requests
- Edit the deployment:
kubectl edit deployment myapp. - Set
resources.requests.cputo match observed usage, e.g., 500m. - Verify the new pods start with the updated requests.
- Edit the deployment:
- Restart metrics server (if unhealthy)
- Run:
kubectl rollout restart deployment metrics-server -n kube-system. - Confirm the pods are running and the metrics endpoint is responsive.
- Run:
- Enable or correct cluster autoscaler
- If disabled, enable it via your cloud provider’s console or
kubectl apply -f cluster-autoscaler.yaml. - Ensure node labels (e.g.,
node.kubernetes.io/instance-type=standard) match the autoscaler’s configuration.
- If disabled, enable it via your cloud provider’s console or
- Fine‑tune HPA behavior
- Use
behaviorto add stabilization windows or cooldown periods. - Example: add a 60‑second
stabilizationWindowSecondsto slow down rapid scaling.
- Use
Escalation Criteria
If, after applying the above fixes, the HPA still exhibits thrashing, consider the following escalation steps:
- Contact the cluster‑ops team to review node pool sizing and autoscaler configuration.
- Engage application owners to confirm that the workload’s CPU profile has not changed (e.g., new feature release).
- Open a support ticket with the cloud provider if metrics server or autoscaler behavior is inconsistent.
Limitations & Verification Checklist
- Changing
targetCPUUtilizationPercentagecan affect other workloads sharing the same node pool; monitor cluster resource usage after the change. - Increasing pod requests may lead to fewer pods fitting on a node, potentially triggering the autoscaler to add nodes; verify node counts with
kubectl get nodes. - Metrics server must be running version 0.5.0 or newer to support HPA v2beta2; older versions will ignore the
behaviorfield. - After each change, run
kubectl describe hpa <name>to confirm that the event history stabilises.
Practical Check
Once the fixes are in place, monitor the HPA over a full deployment cycle:
# Check HPA events for the last 15 minutes
kubectl describe hpa myapp-hpa | grep -A5 "Events"
# Verify pod CPU usage vs. requests
kubectl top pod -n default | grep myapp
If the event log shows no new SCALE_UP/SCALE_DOWN events and pod CPU usage remains within the requested range, the thrashing issue is resolved.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.