Downsampling vs. Rasterization for Large-Scale Dataset Rendering
29.5K reputation · 22 Aug 2022, 16:51 UTC
When rendering datasets with millions of points in Matplotlib (version 3.0+), the primary constraint is balancing render speed and final file size against the preservation of visual fidelity. Two documented strategies exist to bound the rendering cost for vector backends like PDF or SVG.
Data Reduction
Downsampling via array slicing (e.g., x[::k]) or binning (e.g., hexbin) reduces the number of artists processed. While this minimizes draw time and file size, it risks omitting narrow spikes or critical extrema in the data.
Artist Rasterization
Setting rasterized=True on dense artists allows the backend to render only the data points as a bitmap while keeping axes, labels, and legends as vectors. This preserves every data point in the source but introduces pixelation upon zooming in the output file.
The decision between these methods depends on whether the priority is geometric precision or the avoidance of data loss, yet there is no documented threshold for when one approach becomes more efficient than the other.
- At what point does the overhead of rasterizing a dense artist exceed the cost of downsampling the input array?
- How does the
path.simplify_thresholdinteract with these methods when rendering complex line plots?