Critical health state propagation delay in Consul DNS
25K reputation · 06 Dec 2021, 00:47 UTC
When migrating a small application to new infrastructure using Consul for service discovery, the goal is to achieve zero downtime by leveraging health checks to steer traffic. In this architecture, the Consul DNS interface is used to resolve healthy service instances.
A potential conflict arises when a service instance is marked as critical. While the Consul catalog updates rapidly via the gossip protocol, the propagation of this state to clients relying on DNS can be inconsistent due to caching layers at the OS or application level.
If the TTL of the DNS record is not aligned with the health check frequency, clients may continue to route traffic to a failing or decommissioned node despite the service being marked critical in the Consul UI.
- How does Consul manage the TTL of DNS responses to minimize stale records during a service transition?
- What is the recommended configuration to ensure DNS clients react to a
criticalhealth state without introducing excessive query overhead?