Scale Your Stateless Apps on Kubernetes with CPU‑Based Horizontal Pod Autoscaler
Learn how to let Kubernetes automatically scale your stateless workloads with the Horizontal Pod Autoscaler. Step‑by‑step guide, example, trade‑offs, and checklist to get CPU‑based autoscaling working safely.
03 Sept 2026, 22:26 UTC

Why CPU‑Based Autoscaling Matters
Stateless web services often see traffic that spikes for minutes and then drops back to a low baseline. If you over‑provision replicas, you waste compute and cost. If you under‑provision, users experience timeouts. The Horizontal Pod Autoscaler (HPA) lets you let Kubernetes do the balancing: it watches real CPU usage and adjusts replicas in the Deployment automatically.
How HPA Works Internally
The HPA is a control loop that runs inside the kube-controller-manager. It periodically (default every 15 s) pulls metrics from the metrics-server API. Those metrics are compared to a target utilization you set (e.g., 50 % CPU). If the observed value is above the target, the HPA increases replicas; if below, it decreases them, respecting minReplicas and maxReplicas.
Key components:
- Metrics API:
/apis/metrics.k8s.io/v1beta1provides per‑pod CPU/memory usage. - Control loop: Runs every 15 s by default; you can tweak with
--horizontal-pod-autoscaler-sync-period. - Target utilization: Declarative; set in the HPA manifest.
Prerequisites: Install the Metrics Server
HPA depends on a functioning metrics-server. If you’re on a managed cluster (EKS, GKE, AKS) it may already be present. Otherwise deploy it with the official manifests:
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
Verify:
kubectl get deployment metrics-server -n kube-system
kubectl get --raw "/apis/metrics.k8s.io/v1beta1/pods" | head
If the second command returns 404 Not Found or the deployment pods aren’t ready, the HPA will stay in Unknown state.
Deploy a Sample Stateless App
Below is a minimal nginx Deployment that will be autoscaled. It exposes port 80 and uses a simple 2‑byte response to keep CPU usage low under normal load.
apiVersion: apps/v1
kind: Deployment
metadata:
name: nginx-demo
spec:
replicas: 2
selector:
matchLabels:
app: nginx-demo
template:
metadata:
labels:
app: nginx-demo
spec:
containers:
- name: nginx
image: nginx:latest
ports:
- containerPort: 80
Apply it:
kubectl apply -f nginx-deployment.yaml
Create the HPA Manifest
The HPA targets 50 % CPU utilization, with a minimum of 2 and a maximum of 10 replicas. Adjust the bounds to fit your capacity constraints.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: nginx-demo-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: nginx-demo
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 50
Apply the HPA:
kubectl apply -f nginx-hpa.yaml
Verify the HPA is Active
Run:
kubectl get hpa nginx-demo-hpa -w
Initially you’ll see TARGETS as <unknown> until the metrics-server reports data. After a few seconds it should show something like 50%/50% (desired/utilized) and 2/2 for REPLICAS.
Generate Load and Observe Scaling
Use a lightweight HTTP load‑generator such as hey or wrk. Example with hey (install from https://github.com/rakyll/hey):
# Start a burst of 500 requests per second for 60 seconds
hey -z 60s -q 500 -c 10 http://$(kubectl get svc nginx-demo -o jsonpath='{.status.loadBalancer.ingress[0].ip}')
While the load runs, watch the HPA:
kubectl get hpa nginx-demo-hpa -w
You should see REPLICAS climb from 2 to 8 (or your maxReplicas if the load is heavier). After the load stops, the HPA will gradually scale back down, respecting minReplicas.
Trade‑Offs and Limitations
- Latency of Response: The HPA checks every 15 s by default, so rapid spikes may not be fully absorbed immediately.
- Metric Lag: CPU metrics are averages over a 15‑s window; sudden short spikes may be missed.
- Thrashing: If
averageUtilizationis set too low, the system can oscillate between high and low replica counts. Use a safety margin or enablebehaviorto fine‑tune scaling thresholds. - Non‑CPU Workloads: HPA v1 only considers CPU and memory. For I/O‑bound or custom metrics, switch to HPA v2 and deploy a Prometheus Adapter.
Actionable Checklist
- Confirm
metrics-serveris running and healthy. - Deploy your stateless workload with a reasonable
replicasbaseline. - Create an HPA manifest targeting CPU (or memory for HPA v2).
- Apply the HPA and watch
kubectl get hpa -wfor status. - Generate realistic traffic and verify scaling up.
- Stop traffic and confirm graceful scale‑down.
- Adjust
minReplicas,maxReplicas, oraverageUtilizationas needed. - Optional: enable
behaviorto tune scale‑up and scale‑down thresholds. - Monitor
kubectl describe hpafor any warnings or errors.
By following these steps you’ll have a declarative, self‑healing system that keeps cost low during quiet periods and performance high when traffic spikes.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.