Azure Kubernetes Service (AKS)
The AKS SLA only ever covers the Kubernetes API server, never your workloads. There is no native multi-region failover for a cluster — cross-region resilience is a fully custom architecture.
Free tier / single-zone node pool
Free pricing tier and/or node pools without Availability Zone spread. Recommended only for dev/test.
- Replication
- None
- RPO
- N/A
- RTO
- N/A — best-effort only
- Failover trigger
- N/A
- Relative cost
- No extra cost
Free tier control plane costs nothing (node costs are separate either way).
Free tier has no financially backed SLA at all, for any failure scenario.
Shows Unknown at every tier — there's no fixed RPO/RTO for this option; it depends entirely on an architecture you'd have to design and build.
Standard/Premium tier with zone-spread node pools
Control plane and node pools spread across Availability Zones. Kubernetes self-healing (pod rescheduling, node autoscaling) plus AZ placement handles zone loss.
- Replication
- None
- RPO
- 0 for stateless workloads — stateful workloads depend entirely on the backing data tier's own zonal redundancy
- RTO
- Minutes — pod rescheduling and node replacement are automatic but not instant
- Failover trigger
- Automatic
- Relative cost
- Low cost premium
Standard/Premium tier adds a small hourly control-plane charge over Free — zone-spreading the node pool itself costs the same per-node price either way.
The 99.95% SLA covers API server availability only — it says nothing about whether your pods are healthy or serving traffic.
No native feature — custom multi-cluster architecture
AKS has no built-in cross-region failover. The standard pattern is independent clusters per region, configuration synced via GitOps, and traffic steered by Front Door/Traffic Manager. Stateful data replication is entirely the backing store's responsibility.
- Replication
- N/A
- RPO
- Entirely dependent on your architecture and the data stores behind it
- RTO
- Entirely dependent on your architecture — DNS/routing failover plus cluster and application readiness
- Failover trigger
- N/A
- Relative cost
- High cost premium
A second full cluster — control plane and node pools — in another region, kept in sync via GitOps. Roughly double the compute and cluster management overhead.
Shows Unknown at every tier — there's no fixed RPO/RTO for this option; it depends entirely on an architecture you'd have to design and build.
Gotchas
The uptime SLA is about the control plane, not your app
99.95%/99.9% covers Kubernetes API server availability. A perfectly healthy control plane tells you nothing about whether deployments, services, or pods are actually up — that's on your own health checks and readiness probes.
There is no 'AKS regional failover' feature to configure
This is the single most common gap found in AKS resiliency reviews: teams assume multi-region is a setting to enable, when it's actually a from-scratch architecture involving duplicate clusters, GitOps-driven config sync, and a global load balancer — none of which AKS provides out of the box.
Free tier in production is an invisible risk
Nothing stops a Tier 1–3 workload from quietly running on a Free-tier cluster with no financially backed SLA — it looks identical to Standard in the portal until you check the tier setting.