Test your DR plan, don't just configure it
A redundancy feature you've enabled but never actually triggered is a hypothesis about how your system behaves, not a fact. All three incidents on the tier guide — GitLab, Delta, British Airways — had backup or failover infrastructure in place. It failed anyway, because nobody had actually exercised it before the day it mattered.
Two kinds of testing, two different questions
They validate different things, and most mature setups use both — not one instead of the other.
Game days
Planned, scheduled, everyone knows it's happening. Validates the process — who declares the incident, who's notified, whether the runbook is actually followable under pressure. Weak at catching technical surprises, because the team is primed and watching.
Fault injection / chaos experiments
Can run unannounced, against production or pre-production. Validates the mechanism — does the failover actually trigger, does it land inside your RTO, does the application degrade the way you assumed. Weak at catching process/communication gaps, because it's scoped to the system, not the humans.
What Azure Chaos Studio actually does
Chaos Studio is Azure's managed fault-injection service — it deliberately breaks things in a controlled way so you find out how your system behaves before an outage tells you. Preconfigured Scenarios (Compute Zone Down, DNS Outage, forced database failovers, cache flushes, network faults) cover the common outage patterns without you having to hand-assemble individual faults, and each run produces a Scenario report — a structured record you can use for retrospectives, stakeholder updates, or compliance evidence.
Microsoft explicitly lists business continuity testing — validating failover behavior and RTO for disaster-recovery plans — as one of Chaos Studio's core use cases, alongside game days, incident reproduction, and using it as a CI/CD deployment gate to catch resilience regressions before they ship. It also names operational-resilience regulations like the EU's DORA as a reason Scenario reports matter — the same evidentiary need behind Australia's APRA CPS 230, which now requires regulated financial institutions to actually demonstrate business continuity, not just document it.
What to test, and how
| You're claiming resilience against | Test it with | Mechanism |
|---|---|---|
| Zone failure | Compute Zone Down scenario | Azure Chaos Studio |
| Region failover (DR) | Full or partial failover drill, plus a forced regional-dependency fault | Chaos Studio + your own DR runbook, executed for real |
| Database failover | Forced SQL/Cosmos DB failover fault | Azure Chaos Studio |
| DNS / network dependency | DNS Outage scenario, network latency/packet-loss faults | Azure Chaos Studio |
| Backup restore | An actual restore into a test vault — not a config check | Azure Backup restore, exercised on a schedule |
| Communication / escalation process | Tabletop exercise — who declares, who's notified, in what order | Process, not a tool. No Azure feature substitutes for this. |
A cadence to start from
Microsoft doesn't prescribe a specific frequency — this is practitioner guidance, not a citation. Match it to what the tier actually costs you if it's wrong.
| Tier | Suggested cadence |
|---|---|
| Tier 0 / Tier 1 | Quarterly, plus after any change to the architecture it depends on |
| Tier 2 / Tier 3 | Semi-annually |
| Tier 4 | Opportunistically — before a major event, or when the tool tells you it's been a while |
Already know what "resilient" needs to mean here? See HA, DR & Backup — where each one belongs for which mechanism you're actually testing.