Using uWSGI’s Cheaper Subsystem to Save Memory Under Variable Load
Learn how to configure uWSGI’s cheaper algorithm to dynamically adjust worker counts, see a concrete example, and understand the trade‑offs involved.
24 Mar 2026, 18:03 UTC

Problem: Wasted memory during low‑traffic periods
Many Python web apps run behind uWSGI with a fixed number of worker processes. When traffic drops, those idle workers still consume RAM, which can be costly on shared hosts or containers with tight memory limits. Conversely, a sudden traffic spike can overwhelm a small static pool, causing request queuing and latency.
Thesis: The cheaper subsystem lets uWSGI grow and shrink its worker pool automatically, keeping memory usage close to actual demand while still protecting against latency spikes.
How cheaper works
uWSGI maintains two bounds: cheaper (the minimum number of workers) and processes (the maximum). The cheaper-algo decides when to add or remove workers. Two common algorithms are:
busyness– measures the fraction of time each worker spends handling requests versus waiting. When average busyness exceeds a threshold, uWSGI spawns more workers; when it falls below another threshold, it terminates excess workers.spare– tries to keep a configured number of idle workers ready, spawning or killing workers to maintain that spare count.
Key tuning parameters:
cheaper-step– how many workers to add or remove at once.cheaper-busyness-max-requests(for busyness) – the request count window over which busyness is calculated.
Worked example: Configuring a Django app with the busyness algorithm
Assume a Django project deployed via uWSGI Emperor. The goal is to keep at least 2 workers idle at night, but allow up to 8 workers during peak hours.
# /etc/uwsgi/apps-available/myapp.ini
[uwsgi]
chdir = /opt/myapp
module = myapp.wsgi:application
master = true
processes = 8 # maximum workers
cheaper = 2 # minimum workers
cheaper-algo = busyness
cheaper-step = 2 # add/remove two workers per adjustment
cheaper-busyness-max-requests = 50 # evaluate over 50 requests
# optional: limit how low busyness can go before scaling down
cheaper-busyness-min = 10 # % busyness below which we shrink
cheaper-busyness-max = 60 # % busyness above which we grow
To apply the change, restart the Emperor (requires root or sudo):
sudo systemctl restart uwsgi-emperor
After restart, you can observe the worker count in real time:
watch -n 1 "ps -C uwsgi -o pid,ppid,pcpu,pmem,args | grep -v grep"
During a low‑traffic period you should see the process list hover near 2 workers. When you generate load (e.g., with wrk -t4 -c100 -d30s http://myapp/), the count will rise toward 8, then fall back after the load stops.
Trade‑off and limitation: Cold‑start latency
The cheaper subsystem reduces memory by terminating workers, but spawning a new worker involves loading the Python interpreter and application code. If traffic jumps faster than cheaper-step can add workers, requests may queue, causing noticeable latency spikes. This is especially true for large frameworks with heavy import costs.
To mitigate:
- Increase
cheaper-stepso more workers are added at once. - Raise
cheaper-busyness-max-requeststo make the algorithm react faster (but watch for premature scaling). - Consider enabling
lazy-appsorlazymode to defer application loading until the first request, trading a bit of first‑request latency for lower idle memory.
Practical verification
After applying a new config, run a simple load test and watch the process list:
- Start a baseline with no load:
ps -C uwsgi --no-headers | wc -lshould approximatecheaper. - Run a short burst:
ab -n 500 -c 50 http://myapp/. - Immediately after, check the count again; it should be closer to
processes(or at least higher than baseline). - Wait a minute with no traffic and verify the count drops back toward
cheaper.
If the worker count does not move as expected, check the uWSGI log for messages like "cheaper: busyness X%" to see whether the algorithm is evaluating correctly.
Actionable closing
Start with modest values: set cheaper to 25% of processes, cheaper-step to 1, and use the busyness algorithm with a 50‑request window. Monitor memory usage (e.g., via ps -o rss) and latency under realistic traffic patterns. Adjust the thresholds gradually until you observe the worker count tracking load without frequent spikes. Remember that any change to the ini file requires a uWSGI restart, so plan for a brief downtime or use rolling restarts if you run multiple instances.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.