Cosmos DB SDK DNS cache not refreshing after regional failover
0 reputation · 30 Jun 2023, 00:58 UTC
Supported feature
Cosmos DB multi-region accounts use a single DNS endpoint with a 60-second TTL that resolves to the nearest regional gateway. Official SDKs (.NET, Java, Node.js) maintain independent DNS caches separate from the OS resolver.
Unresolved behavior
The SDK does not automatically re-resolve DNS after a service-initiated failover event. Applications must implement retry logic with fresh client instances or rely on the SDK's internal retry policy, which may not guarantee immediate DNS refresh. Certificate validation failures during such failovers are also not consistently logged across SDKs, complicating diagnosis.
Goal
Determine whether the SDK's internal DNS cache honors the service's 60-second TTL during failover, or if the cache TTL is fixed at a longer interval that causes stale IP addresses to persist.
Constraints
- SDK version and language runtime affect DNS caching behavior; release notes must be checked for each.
- Disabling hostname verification for custom domains or private endpoints introduces security risks.
- Private endpoint DNS resolution requires Azure Private DNS Zone configuration; misconfiguration resolves to public IPs.
Questions
- What is the effective DNS cache TTL used by the current SDK version, and does it dynamically respect the 60-second TTL returned by the service?
- During a simulated regional outage, does the SDK re-resolve the endpoint hostname on retry attempts, or does it reuse the cached IP until the process restarts?
- Which diagnostic logging flags surface DNS resolution and TLS handshake details to confirm whether certificate validation failures correlate with stale DNS entries?