Identity — the dependency you can't design around
Every service catalog page in this tool asks you to pick a redundancy option, because Microsoft gives you one to pick. Microsoft Entra ID is different: there is no Local/Zonal/Regional choice, no SKU that buys a different failure domain. It's a single, Microsoft-operated global service — and it sits underneath almost every architecture in this catalog, because almost everything authenticates through it. A resiliency plan that carefully zone-redundant every workload but never writes down "what happens if sign-in itself is unavailable" has a gap in exactly the place it can't afford one.
Why it's not a catalog entry
Cosmos DB, Azure SQL, Storage — for those, Microsoft ships several redundancy SKUs and you buy the one that matches your tier. Microsoft Entra ID ships one architecture, run identically for every tenant on the planet, with no regional-pinning or redundancy-level knob exposed to customers. That's not a gap in this tool's coverage — it's an accurate reflection of the product. The thing you actually control isn't which redundancy Entra ID gets; it's how your applications and sessions behave when Entra ID has a bad day.
How Microsoft actually runs it
Entra ID is architected as a global, geographically distributed service from the start — datacenters and regions aren't something a customer opts into, they're baked into how every tenant is served:
"Microsoft Entra ID is a global cloud identity and access management system that provides critical services such as authentication and authorization to your organization's resources."
— Build resilience in your IAM infrastructure with Microsoft Entra ID, Microsoft Learn
The SLA reflects that this is measured as a single worldwide service, not per-region. Microsoft's committed SLA for Entra ID Premium is 99.99% monthly availability, and — notably — this isn't self-reported after the fact only in a legal document: Microsoft publishes actual monthly attainment numbers going back years. Since April 2025 that calculation also counts successful authentications served by the backup system described below, not just the primary path:
What Microsoft actually publishes
| Month | Global SLA attainment | Note |
|---|---|---|
| March 2021 | 99.568% | The March 15–16 key-rotation outage below, in the actual published record. |
| December 2021 | 99.978% | A separate incident month, still within the 99.9% floor but below the 99.99% commitment. |
| Typical month (2022–2026) | 99.998–99.999% | The SLA is met in the overwhelming majority of months — but not literally every one. |
That table — and the methodology behind it — is Microsoft's own public record, not a third-party estimate:
"Performance is measured in a way that reflects customer authentication experience… the number of unique users who both successfully sign in each minute and are issued a token or an error response are tracked as successful user minutes. If a user's sign-in attempt isn't successfully completed within the minute, it's counted as a failed user minute and contributes to a lower availability percentage."
— Service Level Agreement performance for Microsoft Entra ID, Microsoft Learn · SLA commitment figure via SLA for Entra ID (the azure.microsoft.com legal SLA page redirects for automated fetches; this mirrors the same commitment text: "We guarantee at least 99.99% availability of Entra ID," applicable to the Entra ID Premium offering)
What you can actually control
You can't buy a different Entra ID redundancy tier. You can control whether your applications and policies degrade gracefully when the primary path is unavailable — Microsoft ships real mechanisms for this, with real limits worth knowing before you rely on them.
Backup Authentication Service
Microsoft continuously replicates authentication data to a separate system that can issue tokens if the primary service is unavailable. It only covers existing sessions — no new sign-ins, no guest users — but Microsoft notes reauthentications for existing sessions make up over 90% of all Entra authentications, which is why this matters more than it sounds.
Conditional Access resilience defaults
During an outage, policies that can't be fully re-evaluated in real time (group/role membership, sign-in risk, location) fall back to data collected at session start — on by default. You can disable this per-policy for sensitive apps, but Microsoft is explicit that disabling it on a group/role-scoped policy denies everyone in scope, not just the intended users, because membership can't be checked live during the outage.
Hybrid identity has its own failure modes
If you sync from on-prem AD, your authentication path can fail even when Entra ID is perfectly healthy. Password Hash Sync has no on-prem dependency at authentication time; Pass-through Authentication and federation (AD FS) both depend on on-prem agents, servers, and network paths staying up — separate blast radius from anything Microsoft controls.
The backup system's coverage also has real protocol gaps worth knowing before you assume it covers your app: it currently doesn't protect confidential-client web apps using the authorization code flow, the client-credentials flow for service principals, or the resource owner password credentials grant — and native/SPA apps only get covered if they use MSAL correctly and persist tokens for the silent-refresh path.
"The Backup Authentication Service doesn't support new sessions or authentications by guest users."
— Resilience defaults for Microsoft Entra Conditional Access, Microsoft Learn · protocol coverage via Backup Authentication System: Application Guidelines, Microsoft Learn · hybrid failure modes via Build more resilient hybrid authentication in Microsoft Entra ID, Microsoft Learn
This isn't hypothetical
The March 2021 SLA dip in the table above wasn't a rounding artifact — it was one specific, public, well-documented outage.
Azure AD (now Microsoft Entra ID) outage, March 15–16 2021
Starting around 19:00 UTC on March 15, 2021, an automated key-rotation system incorrectly removed a signing key that had been deliberately marked 'retain' to support an in-progress cross-cloud migration — a bug caused the automation to ignore that flag. The missing key broke OpenID Connect token validation broadly. Sign-in failures cascaded into Microsoft 365, Teams, Exchange Online, SharePoint Online, OneDrive for Business, Outlook.com, Xbox Live, Intune, and the Azure Portal itself — not because those services failed, but because none of them could authenticate users. Full resolution took until roughly 09:25 UTC on March 16, about 14.5 hours later, with some services like Intune needing extra recovery time due to caching.
Every one of those downstream services was individually healthy. The single shared authentication dependency being down was enough to take all of them offline at once — which is exactly the blast-radius pattern a redundant, zone-spread architecture doesn't protect you from.
BleepingComputer: Microsoft explains the cause of yesterday's massive service outageSo what do you actually do about it
You can't buy your way to a different Entra ID redundancy tier the way you can with Cosmos DB or Azure SQL. What a resiliency plan can still do:
Write down the blast radius
List every application that authenticates via Entra ID — which is likely nearly everything in this catalog using Microsoft Entra auth. An Entra outage isn't a risk to one service's tier; it's a correlated failure across all of them simultaneously, and your DR runbook should say so explicitly rather than assume it's covered by each service's own redundancy.
Plan for degraded mode, not zero mode
The Backup Authentication Service means already-signed-in users on supported protocols likely keep working during an outage — new sign-ins and guest access likely don't. Know which of your flows fall into each bucket ahead of time, verify your apps actually use MSAL and persist tokens the way the backup system requires, and don't assume "it just always works."
Identity secrets need their own plan too. See Key Vault — another auth-adjacent dependency almost everything else in this catalog quietly relies on.
This is one dependency among several. See HA, DR & Backup for how identity fits into a full resiliency plan, not just data replication.