Scenarios
Serverless API Backend
Azure Functions behind API Management as the public gateway, backed by PostgreSQL Flexible Server, with Managed Disks resiliency behind whatever VM-based tooling supports the pipeline.
Tier 1 — 1 hour RPO/RTO. Business-critical. Significant business impact if unavailable for more than an hour.
Zonal (AZ) redundancy
Bottlenecked by Azure Functions
Regional / cross-region
Bottlenecked by Azure Functions
| Component | Zonal (AZ) redundancy | Regional / cross-region | |
|---|---|---|---|
| Azure Functions Compute | RPO 0 — the compute layer holds no state to lose. If the host storage account uses ZRS (required for this to actually work), Azure Storage replicates that data synchronously across zones too.·RTO Microsoft describes 'brief interruptions that typically last a few seconds' while connections reroute to instances in healthy zones — but this isn't a committed figure. Microsoft explicitly states it does not guarantee that replacement instances for the lost zone's capacity actually get created ('the platform attempts to backfill lost instances on a best-effort basis... doesn't guarantee success'); its own mitigation advice is to overprovision always-ready instances. | RPO Entirely dependent on the backing services (event source, storage, database) behind the function app — Functions itself holds no durable state to lose·RTO Entirely dependent on your own failover architecture and, for the active-passive pattern, on the failover time of whatever event source drives the switch (e.g. Service Bus/Event Hubs geo-disaster recovery) | |
| API Management Integration | RPO Gateway configuration (APIs, policies) typically propagates between zones in under ~10 seconds — not an SLA commitment, just documented typical behavior. The internal cache and rate-limit counters are not fully protected: cache data isn't guaranteed to persist through a zone loss, and rate-limit counters may be out of date on the surviving zones afterward.·RTO No downtime expected for automatic or zone-redundant configurations — the platform detects a zone failure and reroutes traffic to the remaining zones on its own, including for single-unit instances (whose compute resources are already split across two zones). In-flight requests connected to the failed zone are dropped and must be retried by the client. | RPO Gateway configuration typically propagates to every region in under ~10 seconds — not SLA-backed. Rate-limit counters and the internal cache are region-local and are never replicated between regions at all, by design.·RTO No gateway downtime is expected — API Management detects a regional failure and automatically routes traffic to gateways in the surviving regions, which keep serving the most recently synced configuration. But if the primary region itself goes offline, the management plane and developer portal become unavailable and stay that way until the primary region recovers — Microsoft publishes no RTO for that part, and there is no way to push configuration changes or use the developer portal from a secondary region during the outage. | |
| Azure Database for PostgreSQL Flexible Server Data | RPO 0 — Microsoft states zero data loss for both planned and unplanned zone-redundant failover·RTO 60-120 seconds for automatic failover per Microsoft's documented range — but Microsoft's own docs note the failover process 'might take longer than 120 seconds' if the standby has a large recovery backlog to replay before promotion | RPO Not guaranteed. Microsoft's guidance: lag 'typically ranges from a few seconds to minutes,' and 'in some heavy workload or high-latency scenarios, this delay could extend to hours.' A forced promotion during an outage permanently loses any unreplicated data.·RTO The forced-promotion step itself 'typically completes within 1-3 minutes' once triggered — but Microsoft is explicit that you're responsible for detecting the regional outage and triggering promotion; there's no published figure for that detection-and-decision time, so total time-to-recovery isn't bounded | |
| Managed Disks Storage | RPO 0·RTO 0, automatic, if only the disk's own zone is impacted and the attached VM stays healthy — Azure transparently redirects I/O to a replica in a healthy zone. If the VM itself is also down, recovery is manual: force-detach the disk and reattach it to a VM in a healthy zone (data disks only) — unless it's a shared ZRS disk with a standby VM already running in another zone, where SCSI persistent reservation lets the standby take over. | RPO Entirely dependent on which mechanism you configure — ASR quotes as low as ~30 sec–5 min (not SLA-backed); Azure Backup's cross-region restore is only as fresh as the last completed backup job·RTO Entirely dependent on the mechanism you've configured and whether you've actually tested the failover/restore path | |
| Key Vault Security | RPO 0·RTO Near-instant — fully platform-managed, not customer-observable | RPO Not zero, and not committed — replication to the paired region is asynchronous, so changes made shortly before a region failure can be lost. Microsoft publishes no RPO number for this feature.·RTO Not customer-controlled and not committed — Microsoft's own guidance says failover 'can take several hours after the loss of the primary region, or longer,' and isn't necessarily aligned with when other Azure services in the region fail over. Your vault may just be unavailable until Microsoft acts. |