Short answer
In most multi-node Appwrite deployments, latency spikes under heavy writes are not caused by the broker's queue depth alone — they are caused by the realtime service consuming Redis messages slower than writes publish them. The realtime service is a separate process that, by default, runs as a single worker. Once its CPU saturates, Redis pub/sub messages pile up faster than they are fanned out to WebSocket clients, and lag grows linearly with write rate. Broker tuning helps only after you confirm the consumer — not the broker — is the constraint.
Likely explanation vs. confirmed facts
Confirmed by architecture: every database write publishes an event to Redis; the realtime service subscribes, runs permission checks and payload serialization per subscribed client, then pushes over WebSocket. Each of those steps adds per-message CPU work on a process that does not parallelize by default.
Likely, but verify in your deployment:
- Realtime worker CPU saturation (most common cause of growing lag).
- MariaDB lock contention or missing indexes on
documents for your subscription filters, which lengthens each write and therefore the publish interval. - WebSocket reconnect storms from load balancers with short idle timeouts, which look like latency spikes.
Which one dominates depends on your write rate and client count — the diagnostics below separate them.
Diagnose in this order
- Check realtime CPU:
docker stats appwrite-realtime. Sustained usage near one core means worker saturation — this is your answer, skip broker tuning. - Check Redis throughput:
docker exec appwrite-redis redis-cli INFO stats | grep instantaneous_ops_per_sec. If ops/sec plateaus while writes increase, Redis itself is saturated (it is single-threaded); otherwise the consumer is the bottleneck. - Scrape the metrics endpoint (
:3001/metrics on the realtime service) and watch the realtime lag metric over time. A monotonically growing value confirms a consumption backlog rather than transient jitter. - Check the database side: run
EXPLAIN on the queries backing your subscriptions. Full scans (type=ALL) on documents mean each write holds locks longer; add indexes via a migration for the filtered columns (e.g., collection and permission fields). - Rule out the client: reproduce with a minimal
wscat subscriber before changing infrastructure.
Mitigation, scoped to the cause
- Worker saturation: enable realtime scaling (
_APP_REALTIME_SCALING_ENABLED) and raise _APP_REALTIME_WORKERS up to available CPU cores. This requires a shared Redis instance and a WebSocket-aware load balancer (sticky sessions). Do not raise workers beyond cores without also confirming Redis and MariaDB headroom. - DB contention: add the missing indexes and reduce per-write payload size; large documents multiply serialization cost per subscribed client.
- Reconnect storms: raise idle timeouts on the load balancer/overlay network to exceed the client's heartbeat interval.
Broker-level knobs (queue size, prefetch) only matter if step 2 shows Redis itself saturated — a less common case, since pub/sub has no persistent queue; slow consumers simply fall behind. Note also that Appwrite realtime does not guarantee cross-collection ordering, so out-of-order delivery under load can be mistaken for lag.
One open variable
If you can share whether lag grows monotonically with write rate (consumer backlog) or spikes and recovers (reconnects or lock waits), the recommendation narrows to a single fix. Version note: scaling flags described here apply to Appwrite 1.5+; verify against your deployment's .env and release notes before applying.