Skip to main content
Azure Resiliency Map
Choosing your tier

GSLB — public vs. internal

Global server load balancing — routing traffic to the healthiest region — is the traffic-steering half of regional DR. For public-facing services it's a solved problem in Azure. For services that must never be internet-reachable, it structurally isn't, and the gap catches teams who assume "we'll just point Traffic Manager at it" will work the same way it does for anything public.

Public-facing: two real options

Azure Front Door

A global anycast edge service for HTTP(S) only. Automated health probing and failover are built into every configuration — set backend priority and Front Door routes around an unhealthy region on its own, typically well under a minute after a probe fails. Premium SKU can also reach a private origin via Private Link, though the entry point at Front Door's edge stays public regardless.

Azure Traffic Manager

DNS-based and protocol-agnostic — it works for anything, not just HTTP(S), because it only ever hands back an IP address. The tradeoff is that failover is bounded by DNS caching, not a probe interval: clients holding a cached answer keep hitting the dead endpoint until their resolver's TTL expires, so real-world failover is often slower than the health-check timing alone suggests.

Internal-facing: the hard truth

Both of the above are built to probe endpoints that Microsoft's own infrastructure can reach. Neither is designed to reach in — and Microsoft says so explicitly, not as an inferred limitation:

"Azure Traffic Manager health probes are designed to monitor endpoints that are reachable by the Traffic Manager probing infrastructure. Traffic Manager isn't designed to probe endpoints that resolve to addresses within private, non-routable, or Microsoft-internal network spaces. Endpoints that resolve to addresses within these network spaces must be configured as Always serve traffic endpoints. Traffic Manager can't perform health validation for these endpoints and therefore can't provide health-based failover."

— Azure Traffic Manager endpoint monitoring, Microsoft Learn

"Always serve traffic" isn't a degraded mode of health-based failover — it's no failover at all. Traffic Manager stops checking the endpoint entirely and always returns it, healthy or not. Front Door Premium's Private Link feature doesn't change this calculus either: it privatizes the origin, but Front Door's own client-facing entry point is a public anycast address by design — there is no configuration that makes Front Door itself internal-only.

What teams actually build instead

None of these are a managed Azure product — they're patterns you own, build, and maintain, which is the real cost of an internal-only requirement.

Private DNS + health-check automation

A scheduled function or Logic App probes each region's internal endpoint and updates a Private DNS zone record — effectively rebuilding Traffic Manager's DNS-failover logic yourself, privately. Community reference implementations exist (e.g. a DIY private traffic manager pattern on GitHub) — there is no Microsoft-shipped equivalent.

App-level region selection

A regional internal Load Balancer per region, with the calling application deciding which region to use — via a config/service-discovery layer, a circuit breaker, or a retry-with-fallback pattern. Puts the failover logic in code you control and test, at the cost of every caller needing to implement it.

Network/BGP-level failover

Route injection via ExpressRoute/Route Server to shift traffic at the network layer during a regional failure — an advanced, network-team-owned pattern, usually reserved for active/passive scenarios with real operational maturity behind it.

Which one do you actually need?

SituationUseWhy
Public, HTTP(S)Azure Front DoorAnycast edge, automatic health-probe failover, WAF and caching if you want them — the richest option for web traffic.
Public, non-HTTP (TCP/UDP, custom protocol)Azure Traffic ManagerDNS-based, protocol-agnostic — Front Door only understands HTTP(S).
Public frontend, private originFront Door Premium + Private Link to the originKeeps the backend off the public internet; the entry point at Front Door's edge is still public either way.
Internal-only, active/passive is acceptableTraffic Manager, endpoints forced to "Always serve traffic"The only way to point Traffic Manager at a private endpoint at all — but this disables health checking entirely, so "failover" means someone noticing and flipping it by hand.
Internal-only, needs real automatic failoverBuild it yourself — Private DNS + health-check automation, or app-level region selectionNo first-party managed GSLB reaches a truly internal endpoint. See below.

How this stacks up against your tier

GSLB is entirely an RTO mechanism — none of these route on a data-loss basis, so RPO doesn't apply. Same evaluation logic as the rest of this tool: automatic and SLA-backed passes, numerically-fits-but-unverified is a risk, no committed number is unknown.

OptionRTOMeets tier
Azure Front DoorProbe-interval driven, typically well under a minute — SLA-backed
T0
T1
T2
T3
T4
Azure Traffic Manager (public endpoints)Detection at default probe settings (~2 min) plus your configured DNS TTL — not an SLA-backed completion time
T0
T1
T2
T3
T4
Internal-only (DIY pattern)Entirely dependent on the automation/architecture you build — no platform commitment exists
T0
T1
T2
T3
T4

Front Door passes every tier on its own SLA. Traffic Manager numerically fits everywhere too, but without a completion-time SLA behind it — a risk, not a pass, at any tier. The internal DIY pattern is unknown everywhere, on purpose: that's the actual honest answer until you've built and tested it.

This isn't hypothetical

The failure mode that matters most for GSLB isn't a wrong routing decision — it's the GSLB layer itself becoming the single point of failure it exists to remove.

Provider concentration lesson

Dyn DNS attack, Oct 21 2016

A Mirai-botnet DDoS against Dyn — the managed DNS provider behind traffic routing for Twitter, Netflix, Spotify, GitHub, Reddit, PayPal, Etsy, Heroku, and dozens of other major services — took most of them offline for hours across three separate attack waves in a single day. None of those companies' own regional infrastructure failed. Their shared DNS/GSLB provider did.

GSLB routes traffic to your healthiest region — but if the GSLB layer itself is one provider, that provider is now everyone's single point of failure, no matter how many regions you're running behind it.

Krebs on Security: DDoS on Dyn Impacts Twitter, Spotify, Reddit

When Azure-native isn't enough

Two real gaps Azure-native GSLB doesn't close: genuine health-based failover for internal-only endpoints, and not depending on a single provider for the routing layer itself.

Self-hosted appliances (F5 BIG-IP DNS, Citrix ADC)

Deploy as a VM inside your own VNet, so health probes originate from your network, not Microsoft's public probing infrastructure — genuinely reaching private endpoints Traffic Manager structurally can't. This is the real answer when a DIY script isn't enough and the requirement is internal, automatic, and health-based.

Service mesh (HashiCorp Consul, WAN federation)

Purpose-built for internal service discovery and health checking across datacenters/regions, with prepared queries for geographic fallback — a mature, widely-adopted alternative to hand-rolling Private DNS automation for internal-to-internal traffic specifically.

Multi-provider DNS for the public layer

The direct lesson from Dyn: a secondary DNS/GSLB provider (NS1, Cloudflare, AWS Route 53) alongside Azure Traffic Manager or Front Door means one vendor's outage doesn't take global routing down with it — the practical mitigation many companies adopted after 2016.

See the numbers behind this. The Load Balancer / App Gateway / Front Door catalog entry has the sourced RPO/RTO and cost signals for each option.

GSLB only steers traffic. See HA, DR & Backup for what actually keeps the data behind it consistent.