Managing WAL Growth in Realtime Replication
Supabase Realtime manages the trade-off between event delivery and WAL growth by utilizing a logical replication slot. However, PostgreSQL cannot recycle WAL segments until the replication slot has confirmed the processing of those records. In high-churn environments, if the Realtime server cannot consume the change stream as fast as the database generates it, the WAL will grow indefinitely, regardless of whether Row-Level Security (RLS) filters the final output to the client.
The Mechanics of WAL Accumulation
The primary cause of disk pressure in this architecture is replication lag. Because the replication slot tracks the confirmed_flush_lsn (the last position processed by the consumer), any gap between this value and the current WAL LSN results in retained files on disk. RLS is applied at the broadcast layer, meaning the database must still log and stream every single change to the Realtime engine before it is filtered for specific users.
Monitoring and Thresholds
To prevent disk exhaustion, monitor the gap between the current WAL position and the replication slot's flush position. While specific thresholds depend on your total disk capacity, a widening gap during peak churn is a leading indicator of failure.
Use the following query to verify the current lag in bytes:
SELECT slot_name, active, pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn)) AS replication_lag
FROM pg_replication_slots
WHERE slot_name = 'supabase_realtime';
Mitigation Steps for High-Churn Tables
- Identify High-Churn Tables: Determine if a specific table is generating disproportionate WAL volume (e.g., frequent bulk updates).
- Disable Realtime for High-Churn Tables: If real-time updates are not critical for every single row change, remove the table from the
supabase_realtime publication to stop the WAL generation for that specific source.
- Optimize Update Patterns: Replace high-frequency individual row updates with batched operations or move volatile state (like counters or presence) to a dedicated cache or a table not tracked by Realtime.
- Emergency Recovery: If disk space is critical, dropping the replication slot will immediately reclaim WAL space, but this will cause the Realtime service to lose its place in the stream and may require a service reset.
Diagnostic Requirement
To provide a more specific recommendation on whether to scale the instance or optimize the schema, please provide the average replication_lag value (in bytes) during your peak churn periods.