Choosing a Long‑Term Storage Strategy for Prometheus: Remote Write vs Federation vs Thanos
A decision guide that compares remote write, federation, and Thanos sidecar for retaining Prometheus metrics, with a concrete remote‑write configuration example and verification steps.
01 Jul 2026, 18:40 UTC

Decision: How to retain Prometheus metrics beyond the default retention window
When you need to keep metrics for longer than Prometheus's local storage (or you want a global query view), you must pick a strategy that fits your latency, retention, and operational constraints. The three commonly supported options are:
- Remote Write – Prometheus streams samples via HTTP POST to an external time‑series database (TSDB) such as Cortex, Thanos Receive, or a compatible cloud service.
- Federation – A lightweight pull‑based mechanism where one Prometheus server scrapes selected time‑series from other Prometheus instances.
- Thanos sidecar + ruler – Each Prometheus runs a sidecar that uploads blocks to object storage; a Thanos Querier provides a unified view, and the Ruler enables alerting and recording rules across the cluster.
Comparison table
| Feature | Remote Write | Federation | Thanos sidecar |
|---|---|---|---|
| Latency to remote store | Low (seconds) – samples sent as they are scraped | Higher – depends on scrape interval of the federating server (usually 1‑5 min) | Low for upload (background), query latency depends on object storage |
| Retention | Unlimited – dictated by external TSDB | Limited – federation keeps only recent data (default 15‑day retention on the federating server) | Unlimited – object storage retains blocks indefinitely |
| Operational overhead | Requires running and monitoring an external TSDB and ensuring network reliability | Minimal – just another Prometheus scrape target | Higher – sidecar, object storage bucket, Thanos Querier, and optionally Ruler components |
| Query capabilities | Query via the external TSDB's API (may need learning a new query language) | Limited to the metrics you chose to federate; no cross‑cluster joins | Full PromQL across all clusters via Thanos Querier |
| Setup complexity | Moderate – add remote_write block, TLS/auth if needed | Low – add a federate scrape job | High – deploy sidecar, configure bucket, run Querier/Ruler |
Trade‑off summary
If you need the lowest ingestion latency and unlimited retention without changing your query workflow, remote write is the most direct choice. It adds a dependency on an external TSDB but keeps the Prometheus server itself simple. Federation is attractive when you only need a short‑term roll‑up of a few key metrics and want to avoid any extra infrastructure. Thanos provides a global query view and long‑term storage, but it introduces several moving parts and requires an object‑storage bucket.
Concrete implementation: configuring remote write
The following example shows how to ship all scraped metrics to a Thanos Receive endpoint (or any compatible remote write URL) using basic authentication and TLS. Adjust placeholders to match your environment.
- Edit
prometheus.yml– add or modify aremote_writeblock under the global scope:
global:
scrape_interval: 15s
evaluation_interval: 15s
remote_write:
- url: "https://thanos-receive.example.com/api/v1/receive"
# Basic auth – replace with your own credentials or use bearer_token
basic_auth:
username: "prometheus"
password: "${THANOS_RECEIVE_PASSWORD}"
# TLS – set to false if your endpoint uses a self‑signed cert and you have added it to the trust store
tls_config:
insecure_skip_verify: false
# Optional: reduce bandwidth with snappy compression (default)
compression: snappy
# Example relabeling to drop high‑cardinality labels before sending
write_relabel_configs:
- source_labels: ["__name__"]
regex: "node_cpu_seconds_total;.*"
action: keep
- source_labels: ["instance"]
regex: "^.*:9100$"
action: labeldrop
- Validate the configuration before restarting Prometheus:
# Run from the directory containing prometheus.yml promtool check config prometheus.ymlIf the command returns
Success, the syntax is correct.
- Start (or reload) Prometheus – using your process manager, systemd, or Kubernetes rollout.
- Verify that metrics arrive at the remote TSDB:
- For Thanos Receive, you can query via the Thanos Querier:
curl -G "https://thanos-query.example.com/api/v1/query" --data-urlencode "query=up"and check that you see series from your Prometheus instances. - If you are using Cortex, inspect the
/api/v1/status/buildinfoendpoint or use Cortex's UI to confirm ingestion rate. - Check Prometheus's own logs for lines like
remote storage sent batch of X samples in Y msto confirm the remote write loop is active. - Version compatibility – The
remote_writefeature with TLS options likeinsecure_skip_verifyrequires Prometheus ≥ 2.30. Older releases will ignore or reject those fields. - Network reliability – If the remote endpoint is unreachable, Prometheus will retry with exponential backoff and may temporarily buffer samples in memory. Monitor the
prometheus_remote_storage_queued_samplesmetric to detect backpressure. - Label dropping – Aggressive
write_relabel_configscan inadvertently remove needed labels. Test relabeling in a staging environment first. - Authentication secrets – Avoid hard‑coding passwords in the file; use environment variables or Kubernetes secrets as shown.
Limitations and practical checks
Rollback considerations
Adding a remote_write block does not alter existing stored data; it only changes where new samples are sent. To roll back, remove the block (or comment it out), run promtool check config again, and restart Prometheus. No data cleanup is required.
Takeaway
Choose remote write when you need low‑latency, unlimited retention and are willing to operate an external TSDB. Use federation for simple, short‑term roll‑ups with minimal ops. Opt for Thanos sidecar when you require a unified global query view and are prepared to manage additional components. The configuration snippet and verification steps above let you implement remote write safely and confirm it is working in your environment.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.