Tier exceptions — when the label and the built reality don't match
Every tier on this site is presented as one fixed target for a service. Real organizations run into two situations that don't fit that shape: a service whose actual business criticality changes across the year — Tier 2 during a financial close window, Tier 3 the rest of the time — and a service that carries a Tier 1 or Tier 2 label without an owner willing to fund or operate the high availability that label implies. Neither is a data-model problem this tool can fix for you. Both are governance problems worth naming explicitly, because left unnamed, they become the platform team's problem the moment something breaks.
Pattern 1: the tier changes with the calendar
A service is genuinely lower-impact most of the year, then genuinely high-impact for a defined window — month-end close, quarter-end reporting, year-end financial statements, tax filing deadlines, open enrollment, a retail peak-trading period. The business risk really is different in each window, so treating the service as permanently Tier 2 wastes money the rest of the year, and treating it as permanently Tier 3 risks a real incident during the window that matters most.
Pros
- Spend matches actual risk instead of a permanent worst-case assumption
- Forces an honest conversation about when the service really matters, instead of a default label nobody revisits
- Microsoft's own Well-Architected guidance treats this as legitimate: a disaster recovery solution that "excessively satisfies the workload's recovery point and time objectives" outside the window it needs to is waste, not safety margin
Cons
- The transition itself is the new failure mode — someone has to trigger the uplift before the window starts, and remember to scale back down after
- Every configuration you run needs its own testing and validation — you've effectively doubled the DR testing surface, not kept it the same
- Monitoring and alerting thresholds have to know which tier is active right now, or they'll alert on the wrong baseline
- An incident that starts just before or during the transition window is the hardest case to run a clean postmortem on — which configuration was actually live when it happened?
Pattern 2: the label says Tier 1, the budget says Tier 3
A service is classified as a high tier — because of the data it touches, a compliance requirement, or simply because "it's important" — but the owner has no appetite to fund or operate the redundancy that classification implies. Sometimes that's a reasonable call (the service is being retired soon, or the real business impact of downtime turns out to be lower than the label suggests once someone actually asks). Sometimes it's wishful thinking that quietly becomes the platform team's liability.
"If the SLA serves as a business tactic, the organization might intentionally set it to a high value based on the business owner's goals... It's a calculated risk taken by the business team to benefit the business, and the engineering team is aware of the commitment."
— Architecture strategies for defining reliability targets, Microsoft Learn
Notice the load-bearing phrase: the engineering team is aware of the commitment. Microsoft's own guidance treats a gap between the declared target and the built system as acceptable — but only when it's a conscious, named decision both sides can see, not a silent mismatch that surfaces for the first time during an incident review.
Pros
- Avoids building expensive redundancy for a system whose real business impact doesn't justify it
- Keeps a lightweight support model for services that are genuinely low-maintenance in practice
- Makes sense for a system already scheduled for retirement or replacement — investing in HA for something being switched off is pure waste
Cons
- Anyone reading the tier label — a dashboard, an audit, a compliance attestation — reasonably assumes the infrastructure backs it up
- If the label is ever cited externally (a SOC 2 report, a regulatory disclosure, a customer contract), the gap becomes a compliance finding, not just an internal inconsistency
- Without a written decision, the platform team inherits the blame in a postmortem for "failing to meet Tier 1" on a system they were never funded to build to that tier
Who actually owns the risk
Same underlying issue for both patterns: a mismatch between what's declared and what's built. Without an explicit decision on record, the risk doesn't disappear — it just becomes ambiguous about who was supposed to be managing it.
| Without a documented exception | Service / product owner | Platform support team |
|---|---|---|
| Seasonal tier, uplift forgotten | Owns the business impact of an outage during the critical window, but will point at the platform team first if the escalation happened without a clear trigger record. | Has to prove — after the fact — whether they were ever asked to uplift, or whether the request came too late to execute safely. Without a logged trigger, that's a hard case to win. |
| Seasonal tier, downgrade forgotten | Pays the ongoing premium cost of the uplift indefinitely, usually unnoticed until a budget review asks why this line item never came back down. | Runs a permanently higher-maintenance configuration nobody asked them to keep, and looks like over-engineering in the next architecture review even though it was never their call to revert it. |
| Declared tier, unfunded resilience | Carries the reputational and compliance exposure if the label is ever tested — an audit, a regulator, or a real incident — and the infrastructure behind it doesn't hold up. | Inherits the blame in the postmortem for "failing" a target they were never resourced to meet, because there's no record showing the gap was ever raised, only that it was never fixed. |
The fix isn't a better spreadsheet — it's a signed decision
Both patterns resolve the same way: a short, named, time-boxed record that both sides signed, rather than an assumption either side is carrying silently. This is the same shape as a risk-acceptance record in security governance (ISO 27001, PCI-DSS compensating controls) or a documented waiver in NIST's risk management framework — nothing exotic, just applied to resiliency tiers instead of security controls.
What it should record
The service, its declared tier, the actual tier being run (and for seasonal cases, the exact trigger condition and window — "5 business days either side of month-end," not "around close"), and the reason.
Who signs it
The service/product owner accepting the business risk, and the platform lead confirming what's actually deliverable at each state — the exact pairing Microsoft's own guidance calls out: business sets the target, engineering confirms it's real.
When it expires
A review date, not "indefinitely." A seasonal exception should be re-confirmed every cycle, not assumed permanent; an unfunded-tier exception should be revisited on a schedule so "temporary while we plan the migration" doesn't quietly become a two-year status quo.
If you're already using the Platform Capability Catalog internally, this is a natural fit for its existing notes field per platform, rather than a separate system to maintain.
Real-world shape of this problem
Neither pattern is unusual — both are common enough to have standard names in adjacent disciplines, even though most organizations never write down the specific exception in front of them.
Change freezes during financial close
Enterprise IT governance already has a well-established relative of this pattern: many large organizations impose a change freeze on financial systems during month-end/quarter-end/year-end close, precisely because the business impact of an incident during that window is judged far higher than the rest of the month. That's the same underlying judgment as a seasonal tier bump — "this system's criticality has a calendar" — just applied to change risk instead of infrastructure resiliency. The same reasoning extends naturally to raising the tier, not just freezing changes, for that same window.
Retail peak-trading scaling
Scaling up for Black Friday/Cyber Monday and scaling back down afterward is a widely documented, industry-standard pattern for capacity — the same logic applies one level up, to resiliency tier, for any service whose real business criticality is genuinely concentrated in a known window rather than spread evenly across the year.
Not sure what tier a service needs in the first place? See Choosing your tier for the RTO/RPO decision flowcharts this exception logic sits on top of.
How this site decides Pass/Risk/Fail. See Methodology — the evaluation logic assumes one fixed target tier; this guide is what to do when that assumption doesn't hold.