Skip to main content
Azure Resiliency Map

Service catalog

Local / Zonal / Regional redundancy, failover mechanics, SLA, and known gotchas for each service.

Compute

Virtual Machines / VMSS

IaaS compute. Resiliency is entirely a function of how many instances you run and where you place them — a single VM has no failover story at all.

99.9% (single instance, Premium/Ultra disk) → 99.95% (Availability Set) → 99.99% (Availability Zone, 2+ VMs)

Azure Kubernetes Service (AKS)

The AKS SLA only ever covers the Kubernetes API server, never your workloads. There is no native multi-region failover for a cluster — cross-region resilience is a fully custom architecture.

No financial SLA (Free tier) → 99.9% API server (Standard/Premium, no AZ) → 99.95% API server (Standard/Premium with AZ)

App Service

Zone redundancy requires Premium v2/v3+ and a minimum instance count. There is no native regional failover feature — multi-region App Service is Front Door/Traffic Manager plus a second, independently deployed app.

99.95% (Standard+, multi-instance) — the % doesn't change when you turn on zone redundancy

Azure Container Apps

Zone redundancy is a one-time choice made at environment creation — you can't enable or disable it later, and the SLA doesn't change either way. A fully managed, serverless container host with no durable state of its own and no native cross-region failover.

99.95% — identical whether zone redundancy is enabled or not

Azure Functions

Resiliency here is entirely a function of which hosting plan you pick. The (legacy) Consumption plan has no Availability Zone support at all, full stop — the only fix is migrating to Flex Consumption, Premium, or Dedicated. Those three do support zone redundancy, but — unlike Container Apps — it isn't a one-time, creation-only choice on Flex Consumption or Dedicated; it can be toggled on an existing plan (Microsoft's own docs disagree on whether that's also true for Premium). Functions itself is a stateless compute host: its real RPO/RTO comes from the host storage account behind it (must be ZRS to matter) and whatever your function code talks to.

99.95% (Consumption and Flex Consumption/Premium/Dedicated have distinct SLA wording, same headline number)

Data

Azure SQL Database

Zone-redundant HA (Premium/Business Critical/Hyperscale) is synchronous and near-instant. Cross-region continuity via auto-failover groups is asynchronous and, per Microsoft's own guidance, should be triggered by you — not left to the 'automatic' policy.

99.99% (General Purpose zone redundant) / 99.995% (Business Critical zone redundant)

Azure Cosmos DB

The best-positioned data service in this catalog for tight RPO/RTO — but only if you explicitly enable Per-Partition Automatic Failover (PPAF). The default 'service-managed failover' setting most teams reach for can take an hour or more to trigger.

99.99% (single/multi-region writes) → 99.999% (multi-region multi-write, reads)

Azure Cache for Redis

Every SKU of this service (Basic, Standard, Premium) is retiring on September 30, 2028, in favor of Azure Managed Redis — any resiliency plan built today should target the replacement service. Geo-replication here has no guaranteed recovery point.

None (Basic) / 99.9% (Standard) / 99.95% (Premium, zone redundant)

Azure Databricks

Control plane and compute both spread across zones automatically, but losing a cluster's driver node restarts the whole cluster and its running job — zone redundancy doesn't remove that. Classic and Hybrid are the same workspace type under two different names; Serverless is a genuinely different one, with its own compute plane, its own storage, and no VNet for you to get wrong. No native multi-region capability; Databricks' own managed disaster recovery is a distinct, opt-in feature, not something every workspace gets.

99.95%

Azure Database for PostgreSQL Flexible Server

Zone-redundant HA is synchronous with zero data loss and automatic failover in 60-120 seconds — but same-zone HA, a separately-named and identically-priced option, offers the same automatic failover mechanics while never leaving the zone, so it doesn't survive a zone outage at all. Cross-region continuity is read replicas only: asynchronous, with unbounded replication lag, and Microsoft is explicit that promotion is always customer-triggered, even during a declared regional outage. There is no auto-failover-group equivalent for PostgreSQL Flexible Server.

99.99% (zone-redundant HA) / 99.95% (same-zone HA) / 99.9% (no HA)

Networking

Load Balancer / Application Gateway / Front Door

Load Balancer and Application Gateway are regional services — zone resiliency only. Azure Front Door is the actual cross-region failover mechanism for HTTP(S) traffic, via automated health-probe-based routing.

99.99% (Standard Load Balancer, zone redundant) / 99.95% (Application Gateway v2) / 99.99% (Front Door)

VPN Gateway / ExpressRoute

Zone-redundant gateway VMs are automatic on eligible SKUs — Azure explicitly states you don't need to initiate or validate zone failover. A gateway is always a single-region resource, though: regional DR means deploying independent gateways yourself.

Higher SLA on zone-redundant SKUs; Basic SKU excluded and dev/test only

Azure Firewall

New deployments are zone-redundant by default and automatic — Azure spreads instances across at least two Availability Zones with no configuration needed. A single-zone ("zonal") deployment is now the exception, reachable only through API-based tools. Cross-region DR is always a separate, independently managed firewall.

99.95% (single zone / zonal) → 99.99% (zone-redundant, 2+ AZs)

NAT Gateway

The StandardV2 SKU is zone-redundant by default; the older Standard SKU is a zonal resource pinned to one Availability Zone (or, if you skip zone selection, a nonzonal placement Azure picks for you — the weakest option). The zone configuration is locked in at creation and can't be changed afterward. Like other networking appliances, it's single-region with no native cross-region failover.

SLA credits below 99.99% uptime (2+ healthy VMs required; SNAT exhaustion excluded)

Network Security Group (NSG)

An NSG's rule set is control-plane configuration, not a running instance — it's synchronously replicated across every zone in the region automatically, with nothing to configure or verify. Microsoft doesn't publish a dedicated SLA for it, the same way it doesn't for Virtual Network itself. Cross-region means an entirely separate NSG with its own rules, kept in sync by you.

No dedicated SLA (Virtual Network, which NSGs belong to, has none published)