Architecture Note: Azure SQL Database Auto-Failover Groups for Cross‑Region HA
Learn the requirements, minimal design, trust boundaries, operational checks, failure modes, and design change triggers for Azure SQL Database Auto‑Failover Groups.
17 Oct 2025, 21:43 UTC

When an Azure region suffers an outage, applications that depend on a single Azure SQL Database can experience downtime until manual steps are taken. Azure SQL Database Auto‑Failover Groups provide a built‑in listener that automatically redirects connections to the secondary replica, reducing recovery time without requiring code changes.
Requirements
- Regional separation: Primary and secondary logical servers must be in different Azure regions.
- Same Azure AD tenant: Both servers belong to the same Entra ID tenant; cross‑tenant replication is not supported.
- Matching service tier and compute size: The secondary database must use the same service tier (General Purpose, Business Critical, Hyperscale) and the same hardware generation as the primary.
- Backup retention parity: Backup retention policies must be identical on both sides to guarantee a consistent restore point.
- Read‑only secondary: The secondary is configured as read‑only; write workloads must go to the primary.
Smallest Suitable Design
- Primary logical server: Hosts the read‑write database in region A (e.g., East US).
- Secondary logical server: Located in a different region (e.g., West Europe) and hosts a read‑only copy of the same database.
- Auto‑failover group: Links the two databases, provides a DNS listener (
mygroup.database.windows.net), and enforces the failover policy (Automatic with priority 1 for the primary).
The application connects to the listener name instead of the server‑specific FQDN. When a failover occurs, Azure updates the DNS record of the listener to point to the secondary server’s IP address, allowing the client to reconnect using the same connection string.
Trust and Data Boundaries
Data is encrypted at rest with Azure‑managed keys by default; customer‑managed keys can be added via Azure Key Vault. All inter‑replica traffic travels over Azure’s backbone and is protected by TLS. Unless you explicitly route traffic through a user‑controlled virtual network, the data never leaves the Azure region boundaries.
Operational Checks
- FailoverGroupHealth: Azure Monitor metric that shows Healthy or Critical. A Critical state indicates a synchronization break or network partition.
- FailoverGroupLatency: Measures the lag between primary and secondary; high latency raises the possible data‑loss window (RPO) during an unplanned failover.
- ReadOnlySecondaryLatency: Useful if you offload reporting workloads to the secondary.
To verify the secondary manually, run a read‑only query against the secondary endpoint:
-- Execute on the secondary to confirm read‑only status
SELECT DB_NAME() AS database_name, DATABASEPROPERTYEX(DB_NAME(),'Updateability') AS updateability;
The result should show READ_ONLY for updateability.
Set up alerts in Azure Monitor for FailoverGroupHealth = Critical and for FailoverGroupLatency exceeding your acceptable threshold (e.g., 200 ms).
Failure Modes
- Network partition: If the secondary loses connectivity to the primary, Azure waits for the configured grace period (default 1 hour) before triggering an automatic failover to avoid split‑brain scenarios.
- Regional outage: When the primary region becomes unreachable, the failover policy promotes the secondary after the grace period.
- Secondary deletion or mis‑configuration: Removing the secondary logical server or changing its service tier breaks the group; automatic failover will fail until a new secondary is linked.
Applications must be able to handle transient connection errors and retry logic. Using the Failover Partner parameter in the connection string (or relying on the listener DNS) ensures the driver reconnects to the new primary after a failover.
When to Change the Design
- Latency exceeds 100 ms: Consider adding a read‑scale‑out region or using Azure SQL read replicas to reduce replication lag.
- Data‑sovereignty requires same‑country deployment: Choose a secondary region within the same country or use a local secondary.
- High write throughput: Evaluate sharding, Azure SQL Managed Instance with active‑passive replication, or Hyperscale secondary replicas.
- Need for multiple secondaries: The current model allows only one secondary per logical server; for additional read‑scale you must create separate logical servers and failover groups.
Verification Steps (non‑production)
- Create a primary database in East US on a General Purpose server.
- Create a secondary logical server in West Europe, then add the database to an auto‑failover group with Automatic failover and priority 1.
- Confirm the secondary is ready and read‑only by running the SELECT query shown above.
- Monitor FailoverGroupHealth and FailoverGroupLatency in Azure Monitor; set an alert for Critical health.
- Perform a manual failover from the Azure portal and observe that the listener DNS updates and the application reconnects (you can simulate by stopping the primary server temporarily).
- To test automatic failover, block network traffic from the secondary to the primary (e.g., via a temporary network security group rule) and verify that after the grace period the secondary becomes primary.
Remember that any test that alters the primary’s availability will affect connected clients; perform these steps in a staging environment.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.