Health-Aware Service Discovery with Consul DNS
Learn how to use Consul's health-aware DNS to prevent routing traffic to failing service instances, including configuration examples and DNS caching pitfalls.
23 Mar 2026, 14:15 UTC

The Problem: Routing Traffic to Dead Instances
In dynamic environments, services frequently fail, crash, or become unresponsive. Traditional static load balancing or hard-coded IP lists lead to “black-holing” traffic—sending requests to an instance that is technically online (network reachable) but functionally broken (application crashed). The goal is to ensure that a client resolving a service name receives only the IP addresses of instances currently passing their health checks.
The Solution: Consul DNS and Health-Based Filtering
Consul solves this by combining a distributed service catalog with an integrated health-checking mechanism. When a client queries Consul via DNS, the Consul agent filters the catalog in real-time. If a service instance is marked as critical, it is removed from the DNS response immediately, preventing the client from ever attempting a connection to the failing node.
How Service Registration and Health Checks Work
Consul uses a local agent on every node. This agent is responsible for executing health checks and reporting the status back to the Consul server cluster. The server then updates the global catalog, which the DNS interface queries.
Example Configuration: Registering a Web Service
To register a service with a health check, create a JSON definition file. On a Linux node, place this file in the Consul configuration directory (typically /etc/consul.d/ or a custom path defined by the -config-dir flag).
# /etc/consul.d/web-app.json
{
"service": {
"name": "web-app",
"tags": ["production", "v1"],
"port": 8080,
"check": {
"http": "http://localhost:8080/health",
"interval": "10s",
"timeout": "1s"
}
}
}
Deployment and Verification Steps
- Reload Configuration: Run the following command on the node as a user with permissions to manage the Consul process:
consul reload - Verify Registration: Check that the service is recognized by the cluster:
consul catalog services - Check Health Status: Confirm the instance is passing:
consul health checks web-app - Test DNS Resolution: Use
digornslookupto query the Consul DNS port (default 8600). Run this from a machine that has the Consul agent configured as its DNS resolver:dig @127.0.0.1 -p 8600 web-app.service.consul
Expected Result
The DNS response should return the IP address of the node running the web-app service. If you stop the application on port 8080, the health check will fail, and subsequent dig queries will return an empty result (NXDOMAIN or no A records), effectively removing the node from the rotation.
Critical Limitations and Common Pitfalls
DNS Caching (The TTL Problem)
The most common failure in Consul deployments is over-reliance on DNS caching. Many operating systems and application runtimes (like Java) cache DNS lookups indefinitely or for long periods. If the client caches the IP, it will continue sending traffic to a failed node even after Consul has removed it from the DNS record. To mitigate this, ensure your application's DNS TTL (Time to Live) is set to 0 or a very low value.
Health Check Noise
Setting an interval that is too aggressive (e.g., every 100ms) in a large cluster can lead to “flapping.” This occurs when minor network jitter causes a service to toggle between passing and critical, causing massive churn in DNS records and potentially overloading the Consul servers with state updates.
ACL Permission Blocks
If Access Control Lists (ACLs) are enabled, the Consul agent must have a token with service:write permissions to register the service and service:read permissions for clients to query the DNS. Without these, the service will simply not appear in the catalog, regardless of whether the application is healthy.
Comparison: DNS vs. API Discovery
| Feature | DNS Interface | HTTP API Interface |
|---|---|---|
| Ease of Use | High (Standard tools) | Medium (Requires SDK/HTTP client) |
| Precision | Basic (IP lists) | High (Detailed metadata/tags) |
| Caching Risk | High (OS/JVM Caching) | Low (Direct query) |
| Latency | Very Low | Low |
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.