Scaling Kubernetes Pods with Custom Prometheus Metrics Using Horizontal Pod Autoscaler
Learn how to connect Prometheus metrics to Kubernetes' Horizontal Pod Autoscaler so your pods scale on request rate, queue depth, or any custom signal.
17 Mar 2026, 06:28 UTC

Problem: Scaling on Application‑Specific Signals
Many workloads need to react to business‑level signals such as request rate, queue depth, or latency, not just CPU or memory. The built‑in Horizontal Pod Autoscaler (HPA) can only consume metrics exposed through the Kubernetes Metrics API, which by default provides resource utilization via metrics‑server. To drive scaling from a Prometheus query you must expose that query through the Metrics API using an adapter.
How the Pieces Fit Together
The Prometheus Adapter acts as a bridge: it scrapes Prometheus, evaluates a configured series query, and presents the result as a custom metric via the metrics.k8s.io API. The HPA controller then reads that metric the same way it reads CPU utilization, evaluates it every 15 seconds (default), and adjusts replica count according to the minReplicas, maxReplicas, and scaling behavior you define.
Worked Example: Autoscaling a Web Service by Request Rate
Assume a Deployment named frontend exposes a Prometheus metric http_requests_total. We want to keep the average request rate per pod around 50 req/s.
- Deploy the Prometheus Adapter (official Helm chart) in the
monitoringnamespace, ensuring it has RBAC to readmetrics.k8s.ioand to scrape Prometheus. - Create a adapter rule that exposes
requests_per_second:
# rules.yaml (applied via the adapter's config)
- seriesQuery: 'http_requests_total{job="frontend"}'
resources:
overrides:
namespace: {resource: "namespace"}
pod: {resource: "pod"}
name:
matches: "^(.*)_total$"
as: "${1}_per_second"
metricsQuery: sum(rate(<<.Series>>{<<.LabelMatchers>>}[2m])) by (<<.GroupBy>>)
- Create the HPA that targets the custom metric:
kubectl autoscale deployment frontend \
--cpu-percent=0 \
--min=2 \
--max=10 \
--metric=requests_per_second \
--target-value=50 \
-n production
The resulting object looks like:
apiVersion: autoscaling/v2
kind: HorizontalPodAutcaler
metadata:
name: frontend-hpa
namespace: production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: frontend
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: requests_per_second
target:
type: AverageValue
averageValue: "50"
Trade‑offs and Limitations
Because the adapter must query Prometheus and expose the result via the Metrics API, there is inherent latency: a scrape interval (e.g., 30 s) plus the adapter’s evaluation window can delay the metric seen by the HPA. This can cause brief overshoot or undershoot, especially when traffic spikes sharply. To mitigate, you can:
- Reduce the Prometheus scrape interval for the target metric (trade‑off: higher storage cost).
- Configure the HPA
behaviorstanza to limit rapid scaling (e.g.,stabilizationWindowSecondsfor scaleUp). - Monitor the adapter’s error rate; failed queries appear as
0or stale values, which the HPA treats as no data and will not scale.
The HPA cannot scale to zero for a Deployment using only custom metrics; you would need an external scaler (e.g., KEDA) if zero‑replica idle periods are required.
Verify and Adjust
- Check that the metric is available:
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1/namespaces/production/pods/*/requests_per_second" | jq . - Inspect the HPA status:
kubectl describe hpa frontend-hpa -n productionlooks for the lineCurrent Metricsshowing the observedrequests_per_secondvalue. - Watch replica changes during a load test:
kubectl get deployment frontend -n production -w. - Review adapter logs for any
errorlines:kubectl logs -l app=prometheus-adapter -n monitoring.
If the observed metric consistently deviates from the target, adjust the targetValue or the adapter’s query window (e.g., change the rate(...[2m]) range) to better match your application’s response time.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.