Answer
Consul does not attach a per‑record TTL to its DNS responses; the TTL is governed by the agent’s dns_config block. By default Consul returns a TTL of 0 (no caching), so a change in health state becomes visible to a client as soon as the client’s resolver performs a new query after the health‑check result has been gossiped and the DNS record updated.
Likely explanation
The observable delay between a service being marked critical in the UI and clients stopping to receive its address is bounded by:
- The health‑check execution interval (default
10s).
- The gossip convergence time (typically
200‑300ms per round, converging in 2‑3s for a LAN cluster).
- Any resolver‑side caching introduced by the OS or application (determined by the effective TTL returned by Consul).
If the DNS TTL is higher than 0, the resolver may keep using the stale A/AAAA record until that TTL expires, adding latency equal to the TTL value.
Confirmed facts
- Consul’s gossip pool disseminates health‑check results with a default interval of
200ms, achieving cluster‑wide convergence in roughly 2‑3seconds.
- The health‑check must run and produce a new status before Consul can update the DNS view; thus the health‑check interval is a hard lower bound on propagation delay.
- Consul agent logs contain lines such as
health check: service X changed status to critical and dns: updated service record; the timestamp difference approximates the propagation delay.
Recommended configuration
To make clients react quickly to a critical state without generating excessive query traffic:
- Set the health‑check interval to match your failover target (e.g.,
interval = "5s").
- Configure the DNS layer to return a very low TTL, preferably
0 or 1s, via the agent config:
dns_config {
enable_truncate = false
allow_stale = false
default_ttl = "1s"
}
With default_ttl = "0s" (the default) the resolver queries Consul on every lookup, which adds negligible overhead for low‑QPS services but can increase load under high query rates; a 1s TTL caps the query rate to roughly one request per second per resolver while still keeping stale data under a second.
Verification steps
- Enable debug logging on the Consul agent (
-log-level=debug).
- Force a health‑check change (e.g., stop the service or toggle the check script).
- Note the log timestamp of the health‑check status change.
- Repeatedly query Consul DNS (
dig @127.0.0.1 -p 8600 service.consul) and record when the A record disappears.
- Compute the interval between the two timestamps; it should be close to the health‑check interval plus gossip convergence (
< health‑check interval + ~3s).
- Confirm the effective TTL via
consul info or by inspecting the agent’s dns_config block.
Missing diagnostic detail: The exact health‑check interval and DNS TTL currently applied to the service. Knowing these values determines whether the observed delay is expected or indicates a misconfiguration that would change the recommendation.