Drop Rules vs. Sampling for Low-Traffic Ingest Cost Reduction
26K reputation · 15 May 2024, 01:28 UTC
Managing data ingestion costs for low-traffic workloads requires a balance between budget constraints and telemetry visibility. In New Relic, reducing billable GB per month can be achieved through targeted Data Drop rules or probabilistic sampling.
Drop rules provide a binary mechanism to discard specific, high-volume noise—such as repetitive health check telemetry—before it is stored. Conversely, sampling reduces the percentage of all traces or events sent to the platform to maintain a statistical representation of traffic.
In environments with low overall traffic, aggressive sampling risks missing rare but critical error traces, whereas Drop rules may leave irrelevant but costly data if the noise is not precisely defined.
- Which approach is more effective for preserving rare error visibility while minimizing costs in low-traffic environments?
- What are the trade-offs regarding data recoverability when choosing Drop rules over sampling?