Skip to main content
Azure Resiliency Map
Service catalog
Data

Azure Databricks

Control plane and compute both spread across zones automatically, but losing a cluster's driver node restarts the whole cluster and its running job — zone redundancy doesn't remove that. Classic and Hybrid are the same workspace type under two different names; Serverless is a genuinely different one, with its own compute plane, its own storage, and no VNet for you to get wrong. No native multi-region capability; Databricks' own managed disaster recovery is a distinct, opt-in feature, not something every workspace gets.

SLA 99.95%Last verified 2026-08-22
Local

Single-zone deployment or non-AZ region

In regions without Availability Zone support, or when compute capacity is concentrated in one zone under high demand, the control plane and cluster nodes aren't spread across failure domains.

Replication
None
RPO
0 for control plane metadata; compute-plane cached data is ephemeral and expected to be lost on node failure
RTO
Cluster/job restart time if the driver node is lost — no committed figure for this configuration
Failover trigger
Automatic
Relative cost
No extra cost
Protects against: Node-level failure only — no protection from a zone-wide event

Baseline — you pay for the same VM count regardless of zone placement; this option just doesn't spread them.

Meets tier
T0
T1
T2
T3
T4

Shows Unknown at every tier — there's no fixed RPO/RTO for this option; it depends entirely on an architecture you'd have to design and build.

Zonal

Zone-redundant control plane + compute, ZRS/GZRS workspace storage

Control plane databases replicate synchronously across zones automatically for every workspace type. Below that, zone spread depends on which compute plane is doing the work. Classic compute (Classic/Hybrid workspaces) runs inside your own workspace VNet — Databricks still auto-distributes nodes across zones, but only if your VNet actually has subnets and capacity in each zone, which is on you to provision correctly. Serverless compute (in Serverless workspaces, or run alongside classic compute in a Hybrid workspace) has zone selection handled entirely by Databricks in its own account — no VNet to design or misconfigure. Storage redundancy follows the same split: Classic/Hybrid workspaces expose a customer-visible workspace storage account you can set to ZRS/GZRS instead of the GRS default; Serverless workspaces use Databricks-managed default storage instead, without that same direct SKU choice.

Replication
Synchronous
RPO
0 — control plane expects no data loss during a zone outage; compute-plane cached data is ephemeral and recomputed
RTO
Control plane fails over to healthy zones within about 15 minutes. Losing a cluster's driver node restarts the whole cluster; worker-node loss alone replaces faster.
Failover trigger
Automatic
Relative cost
Low cost premium
Protects against: Datacenter/zone failure

Compute cost is unaffected by zone spread either way. On Classic/Hybrid workspaces, switching workspace storage from the GRS default to ZRS/GZRS changes storage pricing — a modest, storage-only premium.

The 15-minute figure is for the control plane. A lost driver node restarts the whole cluster and its running job — plan job checkpointing accordingly, not just zone redundancy. "Classic" and "Hybrid" are the same workspace type under different names (see gotchas) — the meaningful split for zone resiliency is Classic/Hybrid vs. Serverless compute, not Classic vs. Hybrid.

Meets tier
T0
T1
T2
T3
T4
Regional

Managed disaster recovery (gated), or a customer-built secondary workspace

Databricks is a single-region service with no built-in multi-region capability. Managed disaster recovery is a distinct, gated feature — you apply for access through your Azure Databricks account team before you can use it. Once enabled, it continuously replicates Unity Catalog managed tables (Delta Lake, with data), external tables/volumes (metadata only, not the underlying data), views, functions, and permission grants to a secondary workspace, plus — optionally, off by default — workspace assets like notebooks, jobs, SQL warehouses, clusters, and their ACLs. It also offers an optional stable URL that always resolves to whichever workspace is currently primary. It requires the Premium plan, a separately-billed Mission Critical add-on, and serverless compute enabled on both workspaces (the secondary's serverless compute is what actually reads from the primary's storage during replication). Without it, or for what it doesn't cover, the standard pattern is a second workspace you build and sync yourself.

Replication
Asynchronous
RPO
No fixed number — Databricks doesn't publish a committed RPO even for managed DR. Replication is continuous, but data written after the last synced 'replication point' can be lost during an unplanned failover; you monitor actual RPO trends via a system table
RTO
Failover itself completes in minutes once triggered from the account console — but replicated clusters arrive terminated, SQL warehouses arrive stopped, and job schedules stay paused until you manually resume them
Failover trigger
Manual (customer-triggered)
Relative cost
High cost premium
Protects against: Region-wide outage/disaster — only for what's in scope and actually replicated

A second workspace, its own compute, plus the Mission Critical add-on billed separately on both workspaces for as long as managed DR stays enabled — a real, ongoing cost on top of the duplicate infrastructure.

Managed DR narrows the DIY problem for what it covers, but the exclusion list is long — see gotchas — and you still decide when to fail over. "Minutes" describes the platform action, not the full return to a running, job-scheduled state.

Meets tier
T0
T1
T2
T3
T4

Shows Unknown at every tier — there's no fixed RPO/RTO for this option; it depends entirely on an architecture you'd have to design and build.

Gotchas

high

Databricks has no multi-region capability out of the box

Every workspace is pinned to one region. Even with the managed disaster recovery feature, you're maintaining a second workspace and deciding when to fail over — it narrows the DIY problem, it doesn't remove it.

high

Managed disaster recovery is gated — you can't just turn it on

It isn't a self-service setting. Microsoft's own documentation states: "Managed DR is gated. Apply for access through your Azure Databricks account team." Any DR plan or SLA commitment that assumes managed DR is available the moment you decide you need it is wrong — budget approval and provisioning time into the plan, not just the failover procedure.

high

Managed DR's coverage has real gaps — ML, streaming, and secrets aren't replicated

The exclusion list is long: materialized views, streaming tables, Lakeflow pipelines, the actual data inside managed volumes (only their metadata replicates), Unity Catalog and workspace secrets, ML models, model serving endpoints, vector search indexes, Delta shares, published AI/BI dashboards (only drafts replicate), and Spark Structured Streaming jobs run outside a Lakeflow pipeline are all out of scope. Tables with row filters or column masks, and ABAC-tagged resources, are flagged as "failed to replicate" and block RPO until removed from scope. A workspace built around any of these still needs its own separately-designed DR plan for those specific pieces, even with managed DR enabled and working correctly for everything else.

medium

Setup isn't instant: the Mission Critical add-on is a standing cost, and initial bootstrap can take two weeks

Managed DR requires the Premium plan plus the Mission Critical add-on enabled on both workspaces, billed separately at its own rate on all compute in each workspace for as long as it stays on. For a large workspace, the first replication cycle — the initial bootstrap that copies everything in scope to the secondary — can take up to two weeks before managed DR is even ready to use. Plan the timeline as carefully as the budget.

medium

"Classic" and "Hybrid" are the same workspace type, not two different ones

Microsoft's own documentation states it plainly: "Classic workspaces are referred to as Hybrid workspaces in the Azure portal." There's no separate Classic offering to compare against Hybrid — they're identical, just named differently depending on whether you're reading Databricks docs (Classic) or looking at the Azure portal (Hybrid). The real, resiliency-relevant split is Classic/Hybrid vs. Serverless, which are genuinely different compute planes with different network and storage models.

medium

Classic compute's zone spread depends on your VNet; serverless compute's doesn't

On a Classic/Hybrid workspace, classic compute resources run inside your own workspace VNet — Databricks only distributes cluster nodes across zones if that VNet actually has subnets and capacity in each zone, a customer-owned configuration step that can silently be wrong. Serverless compute (available in Serverless workspaces, or alongside classic compute in a Hybrid workspace) runs in a Databricks-managed compute plane with zone selection and replacement handled entirely by Databricks — there's no VNet to misconfigure in the first place.

medium

A lost driver node restarts the whole cluster, not just a worker

Zone redundancy for compute means nodes get replaced automatically, but if the specific node lost is the driver, the entire cluster and its running job restart. Job checkpointing and idempotent design matter as much as the platform's zone spread.

medium

Classic/Hybrid workspace storage defaults to GRS, not ZRS/GZRS

The workspace storage account Databricks auto-provisions on a Classic/Hybrid workspace starts on geo-redundant storage by default, not zone-redundant — you have to explicitly reconfigure it for zone protection. Serverless workspaces use Databricks-managed default storage instead, without that same customer-facing SKU choice — verify its current redundancy commitment directly with Databricks/Microsoft rather than assuming it matches what you'd pick for a classic workspace.

Sources