Diagnosing Prometheus Scrape Failures in Kubernetes Service Discovery
Prometheus reports down targets for a healthy Kubernetes service this diagnostic guide walks selector mismatches endpointSlice reconciliation and scrape config fixes to restore metric ingestion.
02 Dec 2025, 18:06 UTC

Diagnosing Prometheus Scrape Failures in Kubernetes Service Discovery
Prometheus reports up{kubernetes_name=frontend} 0 for a running Kubernetes service and the Scrape targets UI shows the endpoint as Down with a 404 or timeout error. No new time-series data appears and scrape_series_added_total does not increase. Traffic reaches the service but metrics never enter the pipeline.
In most cases the mismatch stems from how Prometheus discovers scrape targets through the Kubernetes API. The system polls for EndpointSlice objects that map pod IPs to service ports then applies label matching and relabeling rules to determine what to scrape. When those labels or the endpoint mapping are out of sync the target disappears or fails to resolve.
Recognizable condition
- up metric labeled with the service name returns 0
- Scrape targets page lists the target as Down or Unknown with a 404 or timeout status
- Deploying a new service does not trigger a new scrape_series_added_total increment
Cause and diagnostic table
| Cause | Symptom | Check |
|---|---|---|
| Selector label mismatch | targets Down 404timout | Verify prometheus.io/scrape annotation and service selector labels |
| Stale EndpointSlice | 404timout after pod IP change | Run kubectl get endpointslices and compare labels to service |
| Default 1-minute scrape interval | missed metrics from rapidly scaling pods | Check if pod scaling outpaced the interval consider annotation-based relabeling |
| Missing TLS configuration | connection refused when exposing metrics externally | Inspect scrape config for tls_config and secret mounts |
Ordered checks
- Confirm the service and endpoints exist: kubectl get svc <service-name> -n <namespace>
- Inspect EndpointSlices: kubectl get endpointslices --all-namespaces | grep <service-name>
- Verify scrape annotations: Check that the service or pod carries prometheus.io/scrape true or that metric_relabel_configs in prometheus.yml includes the target labels.
- Query the up metric: up{kubernetes_name=<service-name>} in the Prometheus UI or via the API. A value of 1 means the target is actively scraping; 0 or NaN indicates it is not.
- Validate the scrape config syntax: promtool check config --prometheus.url <prometheus-url> prometheus.yml
Fixes tied to findings
Selector mismatch: Add or correct the prometheus.io/scrape annotation on the service or align the service selector labels with those in the Prometheus service_discovery or relabel_configs section. Once the annotation is present Prometheus re-discovers the target within the next service-discovery cycle typically 15-30 seconds.
Stale EndpointSlice: If labels changed after a pod restart or scaling event delete the stale slice or restart the Kubernetes controller manager to force reconciliation. Alternatively add the prometheus.io/port and prometheus.io/path annotations to pin the exact scrape configuration.
Default interval too slow: For workloads that scale rapidly the default 1-minute scrape interval may miss short-lived metric series. Add prometheus.io/scrape_interval 15s as a pod or service annotation. Note that changing the global scrape_interval in prometheus.yml requires a Prometheus restart annotation-based overrides are dynamically reloadable.
Missing TLS: To expose metrics outside the cluster add a tls_config block to the scrape config and mount a TLS secret. A minimal configuration looks like:
scrape_configs: - job_name: kubernetes-services kubernetes_sd_configs: - role: service tls_config: ca_file: /etc/prometheus/tls/ca.crt cert_file: /etc/prometheus/tls/cert.crt key_file: /etc/prometheus/tls/key.key
After the secret is mounted and the config includes tls_config Prometheus can scrape metrics over HTTPS without restarting if only tls and relabel_configs changed; other changes still require a full restart. To reload the Prometheus deployment run:
kubectl rollout restart deployment/prometheus
Escalation criteria
- The service uses a service mesh e.g Istio Linkerd that overrides scrape settings via sidecar injection
- Correcting selectors and endpointSlices does not restore targets after two discovery cycles
- TLS handshake failures persist despite correct secret mounting and tls_config directives
- Large-scale pod autoscaling causes continuous target churn that overwhelms the default discovery loop
Verification after fixes
- Re-run kubectl get endpointslices and confirm labels match the service selector.
- Query the up metric again the target should now show 1.
- Increase scrape_series_added_total via the Prometheus UI or prometheus_api::query to confirm new series are being recorded.
- Run promtool check config prometheus.yml one final time to ensure no syntax errors remain.
When each step is confirmed the Prometheus instance will begin scraping the previously unavailable service and the up metric will reflect healthy ingestion without further manual intervention.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.