Skip to main content
Azure Resiliency Map
Choosing your tier

Test your DR plan, don't just configure it

A redundancy feature you've enabled but never actually triggered is a hypothesis about how your system behaves, not a fact. All three incidents on the tier guide — GitLab, Delta, British Airways — had backup or failover infrastructure in place. It failed anyway, because nobody had actually exercised it before the day it mattered.

Two kinds of testing, two different questions

They validate different things, and most mature setups use both — not one instead of the other.

Game days

Planned, scheduled, everyone knows it's happening. Validates the process — who declares the incident, who's notified, whether the runbook is actually followable under pressure. Weak at catching technical surprises, because the team is primed and watching.

Fault injection / chaos experiments

Can run unannounced, against production or pre-production. Validates the mechanism — does the failover actually trigger, does it land inside your RTO, does the application degrade the way you assumed. Weak at catching process/communication gaps, because it's scoped to the system, not the humans.

What Azure Chaos Studio actually does

Chaos Studio is Azure's managed fault-injection service — it deliberately breaks things in a controlled way so you find out how your system behaves before an outage tells you. Preconfigured Scenarios (Compute Zone Down, DNS Outage, forced database failovers, cache flushes, network faults) cover the common outage patterns without you having to hand-assemble individual faults, and each run produces a Scenario report — a structured record you can use for retrospectives, stakeholder updates, or compliance evidence.

Microsoft explicitly lists business continuity testing — validating failover behavior and RTO for disaster-recovery plans — as one of Chaos Studio's core use cases, alongside game days, incident reproduction, and using it as a CI/CD deployment gate to catch resilience regressions before they ship. It also names operational-resilience regulations like the EU's DORA as a reason Scenario reports matter — the same evidentiary need behind Australia's APRA CPS 230, which now requires regulated financial institutions to actually demonstrate business continuity, not just document it.

Source: What is Azure Chaos Studio? — Microsoft Learn

What to test, and how

You're claiming resilience againstTest it withMechanism
Zone failureCompute Zone Down scenarioAzure Chaos Studio
Region failover (DR)Full or partial failover drill, plus a forced regional-dependency faultChaos Studio + your own DR runbook, executed for real
Database failoverForced SQL/Cosmos DB failover faultAzure Chaos Studio
DNS / network dependencyDNS Outage scenario, network latency/packet-loss faultsAzure Chaos Studio
Backup restoreAn actual restore into a test vault — not a config checkAzure Backup restore, exercised on a schedule
Communication / escalation processTabletop exercise — who declares, who's notified, in what orderProcess, not a tool. No Azure feature substitutes for this.

A cadence to start from

Microsoft doesn't prescribe a specific frequency — this is practitioner guidance, not a citation. Match it to what the tier actually costs you if it's wrong.

TierSuggested cadence
Tier 0 / Tier 1Quarterly, plus after any change to the architecture it depends on
Tier 2 / Tier 3Semi-annually
Tier 4Opportunistically — before a major event, or when the tool tells you it's been a while

Already know what "resilient" needs to mean here? See HA, DR & Backup — where each one belongs for which mechanism you're actually testing.