Scaling Kubernetes Apps on Request Rate with Prometheus Adapter
Learn how to scale Kubernetes workloads based on request‑per‑second using the Prometheus Adapter, avoiding late CPU‑driven scaling and improving latency.
17 Nov 2025, 15:39 UTC

Problem: CPU‑based scaling misses request‑driven latency spikes
Many teams rely on the Horizontal Pod Autoscaler (HPA) with CPU or memory as the sole metric. For web services that are I/O‑bound or experience sudden bursts of requests, CPU utilization can stay low while request queues grow and latency climbs. Scaling only after CPU crosses a threshold means the system reacts too late, leading to degraded user experience.
Takeaway: Use a custom request‑per‑second metric via the Prometheus Adapter so the HPA scales before CPU saturation.
Solution Overview
The Prometheus Adapter acts as a bridge between Prometheus queries and the Kubernetes Metrics API. By exposing a custom metric (e.g., http_requests_per_second) through this bridge, the HPA can make scaling decisions based on application‑specific signals rather than generic resource usage.
Configuring the Prometheus Adapter
First, ensure Prometheus is scraping your application’s metrics endpoint. Then create a ConfigMap that defines how a PromQL query maps to a Kubernetes metric.
apiVersion: v1
kind: ConfigMap
metadata:
name: prometheus-adapter-config
namespace: monitoring
data:
config.yaml: |
rules:
- seriesQuery: 'http_requests_total{job="frontend"}'
resources:
overrides:
namespace: {resource: "namespace"}
pod: {resource: "pod"}
name:
matches: "^http_requests_total$"
as: "http_requests_per_second"
metricsQuery: sum(rate(<<.Series>>{<<.LabelMatchers>>}[1m]))
Deploy the adapter using the official Helm chart, mounting this ConfigMap as --config-map=prometheus-adapter-config. The adapter will then serve the metric at /apis/custom.metrics.k8s.io/v1beta1.
Example HPA Manifest
With the adapter ready, create an HPA that targets the custom metric. The example scales a Deployment named frontend to keep the average request rate per pod below 50 req/s.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: frontend-hpa
namespace: production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: frontend
minReplicas: 2
maxReplicas: 10
metrics:
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: "50"
The HPA will query the adapter for the average http_requests_per_second across all pods of the frontend Deployment and adjust replica count accordingly.
Trade‑offs and Practical Checks
Custom metrics introduce operational complexity: you must maintain a healthy Prometheus stack, ensure the adapter’s queries are performant, and guard against metric volatility. High‑frequency spikes can cause the HPA to flap if the stabilization window is too short.
To verify the setup:
- Confirm the adapter exposes the metric:
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | jq .(look forhttp_requests_per_second). - Check the HPA status:
kubectl describe hpa frontend-hpa -n production– theCurrent Metricsblock should show a value being read. - Simulate load with a tool like
heyorwrkagainst the service and watch pod count:kubectl get deploy frontend -n production -w.
If pod count rises as request rate increases and falls after the load stops, the custom metric is driving scaling. Latency of the metric (typically a few seconds due to the scrape interval) should be accounted for in the HPA’s stabilizationWindowSeconds.
Closing
Scaling on request‑per‑second via the Prometheus Adapter lets you react to the true load signal before CPU saturation, improving latency‑sensitive workloads. The trade‑off is added monitoring‑stack responsibility, which can be mitigated by using managed Prometheus services and keeping adapter configurations version‑controlled. Start with a single service, validate the metric pipeline, then extend the pattern to other latency‑critical components.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.