Short answer
There is no documented thread-count threshold where memory suddenly balloons. What actually happens is subtler: each additional Puma thread increases the number of requests being processed concurrently, and every in-flight request allocates Ruby objects. More threads means more simultaneous allocation churn, which grows the Ruby heap and — critically on MRI with glibc malloc — fragments it. The RSS you see is mostly retained memory, not leaked memory, and it scales with concurrency rather than with any magic thread number.
Why threads amplify apparent bloat
Three mechanisms compound in a single-process, multi-threaded setup:
- Concurrent allocation. With N threads, up to N requests allocate at once. Memory-heavy gems (image processing, large JSON/XML parsing, big CSV exports) multiply their peak footprint by the number of simultaneous executions.
- Heap fragmentation. MRI's GC frees objects but rarely returns pages to the OS, and glibc's per-thread malloc arenas make this worse. RSS climbs during traffic spikes and stays there. This looks like a leak but is usually retained, fragmented heap.
- Thread-scoped state. Code using
Thread.current, per-thread memoization, or shared mutable caches grows with thread count and can hold references indefinitely.
Confirmed vs. assumed
Confirmed, well-established behavior: MRI does not aggressively return memory to the OS, and glibc arena behavior worsens it. What is not established is any universal "threads vs. workers" crossover point — it depends entirely on your app's allocation profile. A worker process has a higher fixed baseline (a full copy of the app, though copy-on-write forking shares much of it), but it can be recycled independently. Threads share one heap, so one memory-hungry code path inflates the whole process.
How to verify which problem you have
Log RSS and GC stats per request via middleware:
class MemLog
def initialize(app); @app = app; end
def call(env)
status, h, b = @app.call(env)
rss = File.read("/proc/self/status")[/VmRSS:\s+(\d+)/, 1]
Rails.logger.info("RSS=#{rss}kB live_slots=#{GC.stat[:heap_live_slots]}")
[status, h, b]
end
end
Interpretation: if heap_live_slots grows monotonically over hours, you have a genuine leak (hunt it with memory_profiler or derailed_benchmarks in staging). If live slots stay flat while RSS climbs, it's fragmentation/retention.
Mitigations, in order of effort
- Set
MALLOC_ARENA_MAX=2 in the environment — this alone often flattens RSS curves significantly under glibc. Compare curves under identical load before and after.
- Cap concurrency realistically: 3–5 threads is a common sweet spot; going higher rarely helps a low-traffic app and multiplies peak allocation.
- Audit memory-heavy endpoints for unbounded caches and
Thread.current usage.
- If bloat persists, use
puma_worker_killer or periodic restarts as a safety valve — but treat restarts as mitigation, not a fix.
Note: this assumes MRI (CRuby). JRuby and TruffleRuby have different GC and memory behavior. If you can share whether RSS grows steadily or only after traffic spikes, that single data point would distinguish leak from fragmentation definitively.