Diagnosing Prometheus Remote Write Failures: Step‑by‑Step Guide
Step‑by‑step diagnostic guide for Prometheus remote_write failures, covering connectivity, TLS, auth, config, and storage‑side checks with fixes and escalation criteria.
20 Feb 2026, 14:48 UTC

Recognizable condition
Prometheus server logs repeatedly contain messages like failed to send batch to remote storage and the metrics remote_write_samples_total or remote_write_success_samples_total show zero or near‑zero successful samples. The remote_write_status gauge (if exposed) stays at 0, indicating a persistent transmission problem.
Cause / diagnostic table
| Potential cause | Typical evidence in logs/metrics | Quick diagnostic clue |
|---|---|---|
| Network connectivity (firewall, DNS, routing) | Connection timeouts, i/o timeout errors | Cannot reach the endpoint from the Prometheus host |
| TLS certificate problems (expired, untrusted CA) | Handshake failure, x509: certificate signed by unknown authority | TLS handshake fails even though host is reachable |
| Authentication/token issues (missing, wrong, expired bearer token) | HTTP 401/403 responses, unauthorized in logs | Endpoint reachable but server rejects the request |
Misconfigured URL or path in remote_write stanza | 404 Not Found, no route to host for wrong host | Endpoint returns 404 or similar despite correct network/TLS |
| Remote storage ingestion limits or schema mismatch | 429 Too Many Requests, 400 Bad Request, storage‑side error logs | Requests arrive but are rejected by the remote system |
Ordered checks
- Verify endpoint reachability
Run from the Prometheus host (requires shell access, typically root or the prometheus user):
# Replace with your remote_write URL REMOTE_URL="https://remote-storage.example.com/api/v1/write" curl -v -o /dev/null -s -w "%{http_code}\n" "$REMOTE_URL"If curl times out or returns a non‑HTTP error (e.g.,
Could not resolve host), note the output. You can also usenc -zvfor TCP connectivity. - Check TLS certificates and authentication headers
If the URL is HTTPS, inspect the certificate chain:
openssl s_client -connect remote-storage.example.com:443 -servername remote-storage.example.comLook for
Verify return code: 0 (ok). Any other value indicates a problem. For bearer token validation, manually add the token to a curl request:curl -v -H "Authorization: Bearer $TOKEN" "$REMOTE_URL" -X POST --data-binary @- <<< ''A 401/403 response points to token issues.
- Inspect the Prometheus configuration file
Locate
prometheus.yml(usually/etc/prometheus/prometheus.yml) and examine theremote_writeblock:remote_write: - url: "https://remote-storage.example.com/api/v1/write" bearer_token_file: /etc/prometheus/token tls_config: ca_file: /etc/prometheus/ca.crt cert_file: /etc/prometheus/client.crt key_file: /etc/prometheus/client.keyEnsure the URL matches the endpoint, the token file exists and is readable, and TLS file paths are correct.
- Review remote storage side logs
Ask the storage team to check logs for incoming requests from the Prometheus host’s IP. Look for:
- Authentication failures
- TLS handshake errors
- 4xx/5xx responses (e.g., 429, 400)
- Schema validation errors
If the storage system exposes its own metrics, verify that request counters increase when Prometheus attempts to send data.
Fixes tied to findings
- Connectivity failure: adjust firewall rules, correct DNS entries, or fix routing. After changes, re‑run the curl test.
- TLS/certificate failure: renew or replace the expired certificate, update the CA bundle, or correct
tls_configpaths. Reload Prometheus (kill -HUP $(pidof prometheus)) after editingprometheus.yml. - Authentication failure: generate a new bearer token, store it in the configured file, and ensure file permissions allow the prometheus user to read it (
chmod 600). Update any other Prometheus instances that share the same token. - Misconfigured URL/path: edit the
url field inremote_writeto the correct endpoint, then reload Prometheus. - Remote storage rejection: work with the storage team to increase ingestion limits, adjust schema validation, or clear temporary backpressure. If the storage returns 429, consider enabling
queue_configwith appropriatemax_samples_per_sendandmax_shardsto smooth bursts.
Escalation criteria
If after applying the relevant fix the error persists for more than 5 minutes or the remote_write_failed_samples_total metric shows >10% of total samples, proceed as follows:
- Engage the remote storage team or vendor support.
- Enable debug logging in Prometheus by setting
--log.level=debugand restarting. - Collect a recent snippet of the Prometheus log (
journalctl -u prometheusor/var/log/prometheus/*.log) and share it with the support team. - Consider temporarily disabling remote_write (
remote_write: []) to protect local disk space while the issue is investigated, ensuring local storage retention is sufficient.
Verification
After a fix, confirm success by:
- Checking Prometheus logs for lines like
Successfully sent batch to remote storage. - Observing the metric
remote_write_samples_totalincrease (e.g.,promtool query remote_write_samples_totalor via the UI). - Querying the gauge
remote_write_status; a value of1indicates the last batch succeeded. - Optionally, push a test metric via the Pushgateway or a short‑lived scrape job and verify it appears in the remote storage after a few seconds.
Limitations and practical checks
Increasing the remote_write queue size indefinitely can lead to out‑of‑memory (OOM) conditions; always monitor process_resident_memory_bytes and keep queue_config within sensible bounds. Changing authentication tokens without updating all Prometheus instances may cause partial data loss; roll out token changes in a coordinated fashion or use a shared secret management system.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.