Enable and Verify Horizontal Pod Autoscaler on a DigitalOcean Managed Kubernetes Cluster
Step‑by‑step guide to deploy an HPA that scales a sample nginx Deployment based on CPU utilization, with checks and rollback options.
18 Sept 2026, 22:19 UTC

Desired Outcome
Configure a Horizontal Pod Autoscaler (HPA) on a DigitalOcean Managed Kubernetes (DOKS) cluster so that the replica count of a Deployment automatically adjusts to keep average CPU utilization near a target value (e.g., 50 %). When load increases, the HPA adds pods; when load decreases, it removes them, up to the defined minimum and maximum replica limits.
Prerequisites
- A running DigitalOcean Managed Kubernetes cluster (Kubernetes v1.23 or newer).
kubectlinstalled locally and configured to access the cluster (typically viadoctl kubernetes cluster kubeconfig save <cluster-name>).- The
metrics-serveraddon enabled. In the DOKS control panel this is toggled under "Monitoring" → "Metrics Server", or it can be installed via Helm if disabled. - Sufficient node pool capacity to schedule additional pods, or the cluster autoscaler enabled to add nodes when needed.
Focused Procedure
-
Deploy a sample application. Create a Deployment and Service for nginx (or any container that exposes a CPU‑intensive endpoint). Save the following manifest as
nginx-app.yamland apply it:
Run:apiVersion: apps/v1 kind: Deployment metadata: name: nginx-deployment labels: app: nginx spec: replicas: 2 selector: matchLabels: app: nginx template: metadata: labels: app: nginx spec: containers: - name: nginx image: nginx:1.25 ports: - containerPort: 80 resources: requests: cpu: "200m" limits: cpu: "500m" --- apiVersion: v1 kind: Service metadata: name: nginx-service spec: selector: app: nginx ports: - protocol: TCP port: 80 targetPort: 80 type: ClusterIPkubectl apply -f nginx-app.yaml(requires cluster admin or edit rights). -
Create the HPA manifest. Define the autoscaler targeting the nginx Deployment, with a minimum of 2 replicas, a maximum of 10, and a target CPU utilization of 50 %. Save as
nginx-hpa.yaml:
Apply:apiVersion: autoscaling/v2 kind: HorizontalPodAutoscaler metadata: name: nginx-hpa spec: scaleTargetRef: apiVersion: apps/v1 kind: Deployment name: nginx-deployment minReplicas: 2 maxReplicas: 10 metrics: - type: Resource resource: name: cpu target: type: Utilization averageUtilization: 50kubectl apply -f nginx-hpa.yaml. -
Optional: ensure resource requests are set. If the Deployment lacks CPU requests, the metrics‑server cannot compute utilization. The example above includes requests; verify with
kubectl get deployment nginx-deployment -o yamllooking forresources.requests.cpu.
Expected Checks
- View HPA status:
kubectl get hpa nginx-hpa. The output should show columns likeNAME,REFERENCE,TARGETS,MINPODS,MAXPODS,REPLICAS,AGE. TheTARGETScolumn should display a value such as50%(or<unknown>/50%if metrics are not yet available). - Describe the HPA for events and current metrics:
kubectl describe hpa nginx-hpa. Look for lines likeCurrent metrics: [ resource cpu on pods: 200m / 500m (40%) ]. - Watch replica changes in real time:
kubectl get hpa nginx-hpa -wwhile generating load (see next step).
Generating Load to Trigger Scaling
To observe the HPA increase replicas, create a temporary busybox pod that consumes CPU:
kubectl run -i --tty load-generator --image=busybox --restart=Never -- \
sh -c 'while true; do :; done'
This pod runs a tight loop that uses CPU. After a few minutes, check the HPA again; you should see REPLICAS increase above the initial 2, up to the maxReplicas limit, as the average CPU utilization across nginx pods rises above 50 %.
When you delete the load generator (kubectl delete pod load-generator), utilization should drop and the HPA will gradually scale back toward the minimum replica count.
Recovery Options
- Temporarily cap scaling. Edit the HPA to set
maxReplicasequal to the current replica count:kubectl patch hpa nginx-hpa -p '{"spec":{"maxReplicas":<current-count>}}'. This stops further upward scaling while you investigate. - Adjust target utilization. If scaling is too aggressive, raise the target (e.g., to 70 %):
kubectl patch hpa nginx-hpa -p '{"spec":{"metrics":[{"type":"Resource","resource":{"name":"cpu","target":{"type":"Utilization","averageUtilization":70}}}]}}'. - Remove the HPA. Delete the autoscaler to revert to static replica count:
kubectl delete hpa nginx-hpa. The Deployment will retain its last replica count unless you manually scale it. - Node‑level issues. If new pods remain in
Pendingstate due to insufficient node capacity, check node events:kubectl get events --field-selector involvedObject.kind=Node. Ensure the node pool has available IP addresses and that the cluster autoscaler (if enabled) can add nodes.
Limitations and Practical Verification
- The HPA relies entirely on the metrics‑server. If metrics‑server is unhealthy,
kubectl get hpawill show<unknown>for current metrics and no scaling occurs. Verify metrics‑server health withkubectl get deployment metrics-server -n kube-systemand inspect its logs:kubectl logs -n kube-system deployment/metrics-server. - HPA only scales based on the metrics you configure. For memory‑based scaling, add a similar
Resourcemetric with namememory. Custom metrics require an adapter (e.g., Prometheus adapter) and are not covered here. - Rapid fluctuations can cause thrashing. To dampen, adjust the
behaviorsection in the HPA spec (stabilization windows, scaling policies). - Always test changes in a non‑production namespace or cluster before applying to critical workloads.
By following the steps above, you can enable, observe, and safely manage the Horizontal Pod Autoscaler on a DigitalOcean Managed Kubernetes cluster, ensuring your workloads respond to CPU load while retaining clear rollback paths.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.