Processes vs Threads Pool for parfor with Large Shared Read-Only Data
0 reputation · 30 Jan 2025, 01:31 UTC
When a parfor loop processes large read-only datasets that exceed per-worker memory limits in a Processes pool, the Threads pool offers shared-memory access but introduces thread-safety uncertainty with third-party libraries. MathWorks documentation recommends Processes for general parfor use and Threads only when workers spend most time in thread-safe MEX or built-in multithreaded functions, yet no heuristic exists to automatically select the pool type based on workload characteristics. Switching pools requires shutting down the active pool and restarting, which disrupts interactive development.
Given a workload where each iteration reads from a 10 GB constant matrix and calls a compiled library with unknown thread-safety guarantees, the trade-off centers on whether memory pressure from process duplication outweighs the risk of silent data corruption in threads.
Which documented criteria should drive the initial pool choice for this memory-sharing versus isolation dilemma? How can thread-safety of the external library be verified without modifying its source code? Does the Parallel Computing Toolbox provide any programmatic way to inspect worker memory layout to confirm shared-data behavior?