HA, DR & Backup — where each one belongs
The most common architecture-review finding isn't a missing redundancy feature — it's a team that believes their zone- or region-redundant setup protects them from something it structurally can't: bad data. Here's the distinction, and where a backup solution — Azure-native or third-party — actually fits.
The confusion, in one picture
A single bad write — a fat-fingered delete, a corrupted deploy, ransomware encrypting a volume — hits your primary. What happens to each layer next isn't the same.
Which one do you actually need?
Match the tool to what actually failed — infrastructure or data. It's a different question than RPO/RTO, and most real incidents are a data problem, not an infrastructure one.
| What failed | What protects you | Why HA/DR alone doesn't | Azure mechanism |
|---|---|---|---|
| Node or disk failure | HA (zone redundancy) | This is exactly the failure class HA is designed for. | Zone-redundant SKU — AZ-spread VMs/AKS nodes, ZRS storage, etc. |
| Full zone / datacenter loss | HA (zone redundancy) | Same layer — the point of spreading across zones. | Zone-redundant SKU, automatic failover between zones. |
| Full region outage | DR (regional failover) | HA doesn't span regions — it was never designed to. | Auto-failover groups, GZRS + account failover, Site Recovery, Front Door. |
| Accidental deletion | Backup | HA/DR replicate a delete exactly as reliably as any other write. | Azure Backup point-in-time restore, soft delete. |
| Data corruption / bad deploy | Backup | Corruption is still a “write” — every replica gets it, correctly and fast. | Azure Backup point-in-time restore to before the bad write. |
| Ransomware encryption | Backup (isolated / immutable) | Encrypted data replicates perfectly. HA/DR just give you encrypted copies faster. | Immutable/soft-delete backup vaults, or an air-gapped third-party copy. |
Azure-native or third-party?
Azure Backup isn't the weak option by default — soft delete is now enforced on every vault (14 days retained free, extendable to 180), specifically to survive accidental or malicious deletion. The real question is what else your estate needs beyond that baseline.
Native (Azure Backup + Site Recovery) is usually enough when
- The estate is pure Azure — VMs, Azure SQL, Files, Blob, AKS.
- Enforced soft delete (14–180 days) covers your deletion/ransomware retention needs.
- The team is already Azure-fluent — another vendor console is overhead, not value.
- Targets sit in Backup/Site Recovery's real range: hours for restore, minutes for orchestrated failover — not sub-minute.
Bring in a third party (Veeam, Commvault, Rubrik, etc.) when
- The estate is hybrid or multi-cloud and needs one console across on-prem + Azure + AWS/GCP.
- You need a copy genuinely outside the blast radius of a compromised Azure tenant — not just soft-deleted inside the subscription an attacker already reached.
- Application-consistent backup for complex systems (SAP, Oracle) goes beyond what Azure Backup natively supports.
- Existing enterprise tooling or compliance reporting is already built around a specific vendor.
Most mature shops run both — Azure-native as the default baseline for everything, plus a third-party or a separately-credentialed Azure subscription specifically for the ransomware-resilience copy. See Azure Backup & Site Recovery in the catalog for the concrete RPO/RTO figures behind each option.
This isn't hypothetical
Two real cases, and the current state of the threat that makes the isolation question urgent rather than theoretical.
Maersk / NotPetya, June 2017
NotPetya wiped roughly 150 of Maersk's domain controllers worldwide within minutes — the same trusted network that made replication fast also let the malware spread through it. The company's data survived only because one domain controller, in a Ghana office, happened to be offline during a power cut and was never touched. An employee flew the drive to another office by hand to begin rebuilding the network. The attack cost Maersk close to $300M.
The thing that made every other domain controller replicate perfectly is the same thing that let the malware reach all of them. The copy that saved the company was the one that wasn't connected.
Wired: The Untold Story of NotPetyaOVHcloud Strasbourg fire, March 10 2021
A fire caused by a faulty UPS destroyed OVHcloud's SBG2 data center in Strasbourg just after midnight and severely damaged the neighboring SBG1 facility, taking millions of customer websites offline. Two customers, Bati Courtage and Bluepad, had paid extra for OVH's backup service specifically because it was contracted to be 'physically isolated' from their production servers — but when SBG2 burned, their backups turned out to be racked in the very same building as the data they were meant to protect, and both were destroyed with it. A French court later ordered OVH to pay them a combined €250,000.
Paying for a backup isn't the same as verifying where it actually lives — 'isolated' is just a line in a contract until someone confirms it's true about physical geography too.
Data Center Dynamics: OVHcloud ordered to pay €250k to two customersBackups are the first target, not an afterthought
Veeam's 2024 ransomware survey found 96% of attacks specifically target backup repositories, and 76% of those attempts succeed at least partially — with 39% of backup repositories wiped out entirely. Sophos's independent research put the targeting rate at 94%.
A backup living in the same trust domain as production isn't a separate layer — attackers now assume it's there and go after it directly.
Veeam: 2024 Ransomware Trends ReportStill working out the target itself? Start with choosing your tier — RPO/RTO first, then this page for what actually delivers it.
Configured it — now prove it works. See Test your DR plan, don't just configure it.
Regional DR needs traffic steering too. See GSLB — public vs. internal.