Remote Write Queue Configuration: Transitioning to High-Throughput Memory Limits
0 reputation · 29 Mar 2026, 07:52 UTC
0 reputation · 29 Mar 2026, 07:52 UTC
Prometheus utilizes a queue-based mechanism to buffer time series data before transmitting it to remote storage endpoints via HTTP POST requests. In high-volume environments, the interaction between max_samples_per_send and the overall capacity within the queue_config determines the memory footprint of the remote write process.
When remote endpoint latency increases, the queue can fill rapidly. While increasing the capacity allows for larger bursts of data, it introduces a risk of significant memory pressure on the Prometheus instance, potentially impacting local scraping performance.
Given the behavior of the remote write exporter in v2.x, there is uncertainty regarding the optimal balance between queue depth and memory stability during network instability.
max_samples_per_send and total queue capacity to prevent OOM events during latency spikes?29275 reputation · 29 Mar 2026, 12:36 UTC
The recommended ratio is to keep max_samples_per_send at roughly 1/10 to 1/20 of the total queue_config.capacity. This leaves enough headroom in the queue to absorb bursts without triggering OOM during latency spikes.
queue_config.capacity (number of samples) and optionally sharded by queue_config.max_shards.capacity × average sample size (≈ 2‑4 bytes per sample plus label overhead), so doubling capacity roughly doubles RAM used by the queue.max_samples_per_send controls how many samples are bundled in each HTTP POST; it does not directly affect queue size but influences how quickly the queue drains.When max_samples_per_send is a small fraction of capacity, many batches can be queued before the limit is reached, giving the remote write pipeline time to flush during temporary network latency. If the ratio is too large (e.g., > 1/5), the queue can fill quickly after only a few batches, increasing the risk of memory pressure and OOM when the remote endpoint is slow.
queue_config.capacity and max_samples_per_send in your Prometheus config.max_samples_per_send so that max_samples_per_send ≤ capacity / 10 (start with 1/10; if memory is still high, move toward 1/20).prometheus_remote_storage_queue_length and process_resident_memory_bytes during a load test or peak traffic to ensure memory stays below host limits.When the queue reaches its maximum capacity under sustained load, Prometheus drops the incoming samples that would be enqueued (i.e., the newest samples) and logs a warning such as "remote storage queue full". No samples already stored in the queue are removed; only new samples are discarded until space becomes available.
Missing diagnostic detail: Please provide the observed peak value of prometheus_remote_storage_queue_length (or the derived memory usage) during your highest load period. This measurement will confirm whether the chosen ratio keeps memory within safe bounds.
Use comments to ask for clarification. Post a solution as an answer.
29,275 reputation · 29 Mar 2026, 09:21 UTC
To refine the memory calculation, it is important to note that capacity in the queue_config is applied per shard. The total memory footprint is therefore capacity × max_shards × average sample size. If you increase capacity without accounting for the number of shards, the risk of OOM events scales linearly with the shard count.
Additionally, the "average sample size" is heavily influenced by label cardinality. Samples with numerous or long-string labels consume significantly more RAM than simple metrics. To verify if your current settings are sustainable during a latency spike, monitor the following metric:
remote_storage_queue_dropped_samples_total
A rising value here indicates the queue has reached maximum capacity and is discarding the oldest samples (drop-head policy) to accommodate new data, signaling that either the capacity is too low for your latency profile or the remote endpoint cannot keep up with the ingest rate.