Scenarios
Microservices on AKS
Containerized services on AKS talking to Cosmos DB, decoupled by Service Bus, with Redis for session/cache state.
Tier 1 — 1 hour RPO/RTO. Business-critical. Significant business impact if unavailable for more than an hour.
Zonal (AZ) redundancy
Bottlenecked by Azure Cache for Redis
Regional / cross-region
Bottlenecked by Azure Kubernetes Service (AKS)
| Component | Zonal (AZ) redundancy | Regional / cross-region | |
|---|---|---|---|
| Azure Kubernetes Service (AKS) Compute | RPO 0 for stateless workloads — stateful workloads depend entirely on the backing data tier's own zonal redundancy·RTO Minutes — pod rescheduling and node replacement are automatic but not instant | RPO Entirely dependent on your architecture and the data stores behind it·RTO Entirely dependent on your architecture — DNS/routing failover plus cluster and application readiness | |
| Azure Cosmos DB Data | RPO 0·RTO Not explicitly quantified by Microsoft — described as fully managed with no required customer action | RPO 0 with Strong consistency; <15 min for Session/Eventual/Consistent-prefix; K-operations/T-seconds for Bounded Staleness (configurable, minimum ~100,000 writes or 300 sec)·RTO Typically ~3 minutes, described as automatic — treated here as guidance rather than a contractual guarantee | |
| Service Bus Messaging | RPO 0·RTO Seconds — automatic | RPO 0 in synchronous mode (at a latency/availability cost); otherwise bounded by your configured max replication lag in async mode·RTO Typically a couple of minutes once triggered — but only ever customer-triggered | |
| Azure Cache for Redis Data | RPO Near-zero but not guaranteed — Redis replication is asynchronous even within a region·RTO Seconds — automatic platform-managed failover | RPO Not guaranteed — Microsoft explicitly states 'no recovery point is guaranteed' for unlinked/failed-over data·RTO A few minutes once manually triggered | |
| Key Vault Security | RPO 0·RTO Near-instant — fully platform-managed, not customer-observable | RPO Not zero, and not committed — replication to the paired region is asynchronous, so changes made shortly before a region failure can be lost. Microsoft publishes no RPO number for this feature.·RTO Not customer-controlled and not committed — Microsoft's own guidance says failover 'can take several hours after the loss of the primary region, or longer,' and isn't necessarily aligned with when other Azure services in the region fail over. Your vault may just be unavailable until Microsoft acts. |