Skip to main content
Azure Resiliency Map

Choosing your tier

Most tiers get picked by gut feel — "this feels important, call it Tier 1." RPO and RTO measure different things and rarely deserve the same answer: RPO is about how much work you can afford to redo, RTO is about how long you can afford to be down. Answer them separately, then combine.

This isn't hypothetical

Four real, publicly documented incidents — one per lesson worth internalizing before you fill in a tier.

RPO lesson

GitLab.com, Jan 2017

An engineer deleted the primary production database. GitLab had five separate backup/replication mechanisms — none had been verified working. ~6 hours of issues, comments, and ~700 new accounts were gone for good; the site was down 18 hours while engineers rebuilt from a stale staging copy.

A backup that's never been tested isn't an RPO — it's a hope.

GitLab's public postmortem
RTO lesson

Delta Air Lines, Aug 2016

A single data-center power-control failure took Delta's reservation systems down for roughly five hours. Backup systems existed but didn't fail over cleanly. ~2,000 flights were cancelled over the following days; Delta disclosed a $150M hit to pre-tax profit.

An RTO target means nothing if the failover behind it was never tested end-to-end at full load.

Delta's own $150M disclosure, via Data Center Knowledge
RTO lesson

British Airways, May 2017

A power surge — from an engineer reconnecting a power supply at a Heathrow-area data center — was severe enough that BA's backup power systems didn't engage either. ~75,000 passengers were stranded over a holiday weekend; owner IAG estimated the cost at roughly £80M ($102M).

Redundant power doesn't help if the failure mode takes out the backup path with it — resiliency has to survive the specific failure, not just a generic one.

IAG's cost disclosure, via Data Center Knowledge
RTO lesson

Azure South Central US, Sept 4 2018

A lightning strike near San Antonio during a severe storm caused a voltage swell that took down cooling in part of the South Central US datacenter; automated safety shutdowns kicked in to protect hardware, but not before storage and networking equipment was damaged. Most services in the region were still down more than 25 hours later, and Visual Studio Team Services (now Azure DevOps) — which stored control-plane data for that region — wasn't fully restored for almost 48 hours, because it also served customers hosted in other regions who depended on that same data. The public status page itself was hosted in the affected datacenter and stayed unreliable for over 12 hours.

A 'regional' outage's real RTO isn't bounded by the region on your architecture diagram — hidden control-plane dependencies, and a status page that goes down with the thing it's supposed to report on, can drag services in supposedly unaffected regions down with it.

The Register: Microsoft reveals train of mistakes that killed Azure in South Central US

This tier framework itself isn't invented for this tool — it mirrors Microsoft's own Well-Architected Framework criticality tiers (Mission Critical → Business Critical → Business Operational → Administrative), which independently land on the same shape: Tier 0 here matches their Mission Critical band — "RTO measured in seconds, RPO approaching zero."

Deciding RTO — how long can it be down?

The driver is business impact, not technical difficulty. A workload that's hard to rebuild but invisible to customers still tolerates a long RTO.

Workload is down right nowDo customers get blocked orrevenue stop within minutes?Would anyone outside the teamnotice within a business day?Is there a manual workaroundthat keeps the business running?Does it block other teams'work, not just visibility?No — internal onlyYes — customer-facingNoYesYesNoNoYesTier 41 weekTier 11 hourTier 015 minutesTier 324 hoursTier 26 hours
A starting point, not a verdict — if your honest answer sits between two branches, that's what Custom is for.

Deciding RPO — how much data can you lose?

The driver is recoverability, not sensitivity. Highly sensitive data that's trivially re-derivable (a cache, a search index) tolerates more loss than boring data with no second copy anywhere.

Data since the last successfulcopy is lostCan it be replayed or re-enteredfrom another system?Regulatory, contractual, orfinancial-integrity requirement?Would replaying it realisticallytake more than a few hours?Would it take more thana full business day?No — only copyYes — replayableNoYesNoYesNoYesTier 11 hourTier 0~0–15 minTier 41 weekTier 324 hoursTier 26 hours
Same caveat — a contractual (not regulatory) requirement, or a replay time you're not confident in, is exactly where Custom earns its place instead of rounding.

Combine them — they rarely match

Run the workload through both flowcharts independently, then take the stricter (lower) of the two results as a floor. If they land more than a tier or two apart, that gap is real information, not an error — it means the honest target has a tight RTO and a loose RPO (or vice versa), which none of the five preset tiers can express on their own.

That's what Custom is for. Set RPO and RTO independently on the tier dashboard instead of rounding to whichever preset is close enough.

The other direction: a service whose RPO and RTO don't match

"Combine them" above is about your own target not fitting one preset tier. The mirror case is a real Azure redundancy option whose RPO and RTO land in different tiers on their own — a genuinely tight RPO paired with a much looser RTO, or vice versa. This isn't hypothetical; it's the actual shape of a real option already in this catalog:

Storage Account — RA-GRS / RA-GZRS with Geo-priority replication (recommended)

RPO ≤15 min, SLA-backed — but only for Block Blobs. Files/Tables/Queues/Page Blobs replicate best-effort with no committed RPO — good enough for Tier 0 on its own. RTO Reads: near-0, automatic, 99.99% SLA on the RA- secondary endpoint. Writes: failover process plus DNS update, typically under an hour once triggered — only good enough for Tier 1. Same option, two different tiers depending on which number you look at.

TierThreshold (RPO = RTO on every preset)This option: RPO 15 min / RTO 1 hrVerdict
Tier 015 minutesRPO fits · RTO too slow
Tier 11 hourRPO fits · RTO fits
Tier 26 hoursRPO fits · RTO fits
Tier 324 hoursRPO fits · RTO fits
Tier 41 weekRPO fits · RTO fits

Notice what happens: it fails Tier 0 outright, even though its RPO alone is a Tier 0 number — the RTO isn't. And where it does numerically fit (Tier 1 and looser), it caps at Risk, never Pass, for a completely separate reason: failover has to be triggered manually, so the RTO clock doesn't start until someone notices and acts. Two independent ceilings, stacked on the same option.

The rule: a service's effective tier is capped by its worse metric, never averaged and never rescued by its better one. Meeting a tier requires both RPO and RTO to fit that tier's thresholds — this evaluator uses an AND, not an OR, deliberately. A business can't say "we only lost 15 minutes of data" as a consolation for a system that was down for a day; both numbers describe real impact, and the worse one is what actually happened to the business during the incident.

The engineering-tooling trap

The git repo, an AI/automation tool, the CI runner, the internal wiki — these get over-tiered more than almost anything else. The person setting the tier is usually the person who feels the outage first and hardest, which is exactly the wrong vantage point for a business-impact question. Proximity to pain isn't proximity to revenue.

Run it through the RTO flowchart honestly: "do customers get blocked or revenue stop within minutes?" For most engineering tooling, the true answer is no — velocity drops, but whatever is already deployed keeps serving customers. That alone routes it to the internal branch, not Tier 1/2. Before you settle there, check three exceptions first:

  • Is it the only copy of something irreplaceable? A git host holding the sole copy of your source history with no mirror is an RPO problem, not an RTO one — assess data loss on that tool separately from its uptime.
  • Is it in the path of fixing other outages? Monitoring, paging, SSO, and whatever CI/CD pipeline ships an emergency fix are force multipliers — their failure cascades into your ability to recover everything else. Tier those closer to your most critical dependent system, not as a generic internal tool.
  • Does it orchestrate a customer-facing process, not just support one? An AI/automation tool running internal reporting is Tier 3/4. The same tool executing order fulfillment or payment reconciliation isn't "an internal tool" anymore — tier it as part of whatever process it drives.

The evidence backs the default, not the instinct: Atlassian's own postmortem on their April 2022 outage — a maintenance script bug that deleted 883 customer sites outright, not a brief blip — reports up to 14 days of downtime for ~775 organizations' Jira and Confluence. That's about as bad as engineering/PM tooling failure gets, and Atlassian's own numbers put it at 0.4% of their customer base, with data loss under five minutes for nearly all of them. Severe for the vendor and painful for the teams involved — but not the kind of impact that justifies defaulting this tool class to Tier 0/1.

Worked examples

Checkout & payments service

  • RTO: customer-facing, blocks revenue immediately, no manual workaround → Tier 0
  • RPO: financial transactions, not replayable, regulatory requirement → Tier 0

Both branches agree — an easy call, and it should feel obviously strict.

Nightly internal BI dashboard

  • RTO: internal only, nobody notices for a business day → Tier 4
  • RPO: fully replayable from the source warehouse, quick to redo → Tier 4

Also agree — resist the urge to round this up to Tier 2 out of caution.

Git repo & CI/CD

  • RTO: engineers can't push or ship, but already-deployed code keeps serving customers and other departments don't notice for a business day → Tier 4, by default
  • Exception: if its CI/CD is the only path to deploy an emergency fix mid-incident, elevate to Tier 2 — the failure cascades into every other recovery
  • RPO: if repos have no external mirror, that's a separate Tier 1 data-loss question about commit history, not an uptime one

The engineering-tooling trap in miniature — the default is low, but two real exceptions can override it.

AI / automation tool

  • If it runs internal reporting/notifications: Tier 4 — same reasoning as any internal tool
  • If it executes a customer-facing process (order sync, payment reconciliation, customer emails): it's not "an internal tool" anymore — tier it as part of that process, which is often Tier 1 or Tier 0

There's no single right answer for an automation tool — only for what it's actually wired to.

Have a tier — now what actually delivers it? See HA, DR & Backup — where each one belongs for which mechanism actually gets you there, and where backup solutions fit.

Have that mechanism configured? See Test your DR plan, don't just configure it — a configured feature you've never triggered is a hypothesis, not a fact.

Multi-region means routing traffic too. See GSLB — public vs. internal — the internal case doesn't work the way most teams assume.

"Regional DR" often means one specific region. See Paired regions, explained — several services' built-in failover only works if your region has one at all.

Every tier here assumes sign-in works. See Identity — the dependency you can't design around — the one dependency with no redundancy tier to pick.

Not every service has one fixed tier, year-round. See Tier exceptions — when the label and the built reality don't match — seasonal criticality and unfunded HA, and who owns the risk when they happen.

Native backup isn't your only option. See Third-party backup & DR — Rubrik, Metallic, Veeam, and Druva compared on what you'd actually be operating, not just feature checkboxes.