Direct answer
There are no universal recommended values for cheaper-busyness-min-requests and cheaper-busyness-max-requests. These thresholds must be calibrated to your request pattern and tolerance for cold-start latency. The busyness algorithm counts requests per worker over its sampling window; min-requests sets the floor below which a worker is considered idle and eligible for termination, while max-requests sets the ceiling that triggers a new worker spawn. Start with a wide gap (e.g., min-requests=10, max-requests=50 over the default 10-second window) and narrow it only after observing oscillation in --stats output.
How the busyness algorithm works
The cheaper subsystem with cheaper-algo=busyness samples each worker's request count every cheaper-busyness-window seconds (default 10s). If a worker's count is ≤ cheaper-busyness-min-requests, it is marked for termination when the pool exceeds cheaper (minimum workers). If the average across all workers exceeds cheaper-busyness-max-requests, a new worker is forked up to processes (maximum). Oscillation occurs when traffic hovers near the thresholds, causing rapid spawn/terminate cycles.
Interaction with lazy-apps = true
When lazy-apps is enabled, the application is imported only on the first request handled by each worker. During a scale-up event triggered by the busyness algorithm, the newly forked worker has no application loaded. The first request to that worker incurs the full import cost (module loading, framework initialization, database pool creation, etc.). This adds cold-start latency on top of the fork time, which can be seconds for large Django/Flask apps.
The busyness algorithm does not account for this extra latency. It sees request counts rising, spawns a worker, but that worker cannot serve requests at full speed until its first request completes the lazy import. If traffic spikes faster than the import completes, the algorithm may spawn additional workers, amplifying memory pressure without immediate throughput gain.
Practical tuning approach
- Measure baseline: Enable
--stats /tmp/uwsgi.stats and run uwsgi --connect-and-read /tmp/uwsgi.stats under realistic low traffic. Note the request-per-worker rate.
- Set initial thresholds: If your idle workers see ~2 req/10s, set
cheaper-busyness-min-requests=5 (terminate below this) and cheaper-busyness-max-requests=30 (spawn above this). The 6× gap dampens oscillation.
- Account for lazy-apps latency: If cold-start import takes >500 ms, increase
cheaper-busyness-max-requests further (e.g., 50–80) so the algorithm spawns earlier, giving the new worker time to warm before the next spike.
- Consider
cheaper-initial: Set cheaper-initial slightly above cheaper (minimum workers) to keep a small warm buffer, reducing scale-up frequency.
- Verify with load test: Use
wrk or hey to simulate a step increase from idle to peak. Watch --stats for worker count stability and latency percentiles.
One missing diagnostic detail
What is the typical cold-start import latency (time from worker fork to first response) for your application with lazy-apps=true? This single number determines whether you need aggressive pre-spawning (higher cheaper-initial) or can tolerate the busyness algorithm's reactive scaling.