Hinted Handoff vs Batch Log: Choosing the Right Strategy for Transient Node Failures in Multi‑Data‑Center Cassandra
0 reputation · 23 Dec 2023, 08:45 UTC
Context
In a single‑node test environment, the default max_hint_window_in_ms of 60 seconds is often sufficient to recover from brief outages. However, in a production cluster spanning multiple data centers, network partitions can last longer than this window, causing hints to expire before they are replayed.
Trade‑Off
Enabling hinted handoff trades increased write latency for eventual consistency, while the batch log persists failed mutations on disk, ensuring durability even after node restarts but at the cost of higher disk I/O and potential write backlog during heavy traffic.
Unresolved Decision
The interaction between hints_ttl, max_hint_window_in_ms, and multi‑data‑center deployments is not fully documented. It is unclear how hint expiration behaves when nodes in different data centers are down for extended periods, and whether disabling hinted handoff during sustained outages has any long‑term impact on cluster health.
Questions
- When should a cluster enable hinted handoff versus relying on the batch log for transient node failures?
- What is the effect of disabling hinted handoff during prolonged outages on eventual consistency and hint expiration across data centers?
- How do
hints_ttlandmax_hint_window_in_msinteract in a multi‑data‑center cluster when nodes experience extended partitions?