mutations_sync = 1 vs Asynchronous Mutations for Schema Consistency
24.5K reputation · 24 Jun 2021, 12:46 UTC
Schema Update Propagation in ReplicatedMergeTree
When executing ALTER TABLE ... UPDATE or similar mutations on a ReplicatedMergeTree table across a cluster, ClickHouse manages these changes asynchronously by default. The metadata is propagated via ZooKeeper, but the actual data rewriting occurs in the background as parts are merged.
In environments requiring strict consistency, there is a trade-off between using the default asynchronous behavior and enabling SET mutations_sync = 1. While the synchronous setting ensures the client session blocks until the mutation is applied, it introduces significant latency and potential timeouts for large datasets.
The uncertainty lies in balancing write availability against the risk of a query hitting a replica where the mutation has not yet completed, leading to inconsistent results across the cluster.
- Does
mutations_sync = 1guarantee that all replicas have physically applied the change before the command returns? - What are the performance implications for insert throughput when synchronous mutations are enforced on high-volume tables?