Manual AG Failover: 5-30 seconds
MultiSubnetFailover configured: Often <10 seconds
Poor client connection handling: 30-120+ seconds
DNS caching issues: Several minutes
The SQL databases themselves are often unavailable for only a few seconds. The majority of downtime seen by users is typically application reconnection time rather than the AG failover itself.
With a planned failover, SQL Server first verifies:
SQL Server first verifies:
Secondary is synchronized
Data movement is healthy
Databases are ready
Then role transition is performed cleanly.
With an unplanned failover or reboot:
WSFC detects node loss. Failover becomes reliant on:
Cluster heartbeat timeouts
Lease timeout detection
Resource restart logic
The AG may not fail over immediately. Applications can sit waiting while the cluster determines the primary is genuinely gone.
Application interruption may only be a few seconds.
The outage window is generally longer because failover only starts after the failure has been detected.
If necessary, action can be taken to suspend workload, drain queues, pause schedulers.
A reboot simply kills everything.
Although, if you plan the reboot, this particular issue can be avoided.(but if you can plan, why choose a reboot over a controlled failover?)Many applications handle a SQL Error better than they handle...
network connection dropped
If you regularly see issues caused by planned failovers, consider improving the application retry logic and scheduler behaviour.
A well designed AG-aware application should normally recover from a planned failover with few or no permanent job failures.There is a choice here... for your application, which has the biggest impact? Some failed transactions? Or a longer outage where all the tasks below are performed before failover? How well your application copes with the failures should guide your decision.
Any application improvements you make now to support planned failovers will pay dividends when you have an unplanned failover.
PERFORM ONLY THE STEPS IN THIS SECTION THAT YOU THINK ARE JUSTIFIED FOR YOUR APPLICATION
Identify if there are long-running transactions in-flight (and choose whether to let them complete or rollback)...
DBCC OPENTRAN;
Stop Job Schedulers...
SQL Server Agent
Windows Scheduler
Stop-Service SQLSERVERAGENT
ALTER AVAILABILITY GROUP [MYAG] FAILOVER;
Must be run from one of the secondary replicas (not primary)ALTER AVAILABILITY GROUP [MYAG] FORCE_FAILOVER_ALLOW_DATA_LOSS;
Must be run from one of the secondary replicas (not primary)TODO: Convert to Image Carousel
During failover:
Listener moves to other node.
Cluster updates ownership.
DNS registration may refresh.
Some applications:
Cache DNS aggressively.
Never retry DNS lookup.
Maintain stale IP addresses.
Result:
Application uses old IP and connection fails
Common with:
.NET
Java
IIS
Application servers
The application keeps a pool of SQL connections.
When failover occurs:
The pool may continue trying to reuse dead connections.
Symptoms:
Login failures
Timeout errors
"Transport-level error"
"Connection reset"
A modern application should detect this and recreate connections automatically.
For modern SQL clients:
MultiSubnetFailover=True
should be used for AG listeners spanning subnets.
Without it:
Client probes IPs slowly.
Failover detection takes longer.
Microsoft specifically recommends this for AG listeners.
Typical effect:
With MultiSubnetFailover - 2-5 seconds
Without - 20-60+ seconds
depending on network configuration.
Any active transaction at failover is terminated and rolled back.
Well-written applications retry the transaction.
Poorly written applications surface an error to the user.
During failover:
Primary jobs stop.
New primary jobs start.
If jobs are not AG-aware you can get:
Duplicate execution.
Missed execution.
ETL failures.
Reporting delays.
An application may connect properly via the Listener but then call:
EXEC ServerA.Database.dbo.Procedure
If ServerA is the Primary all will be fine. If it's not then things will likely break.
If using:
Read-only replicas
Read-intent routing
connections may temporarily fail while:
Replica roles change.
Routing tables update.
Listener redirects clients.
Usually only seconds, but some applications handle the redirect poorly.