Azure Databricks
Control plane and compute both spread across zones automatically, but losing a cluster's driver node restarts the whole cluster and its running job — zone redundancy doesn't remove that. Classic and Hybrid are the same workspace type under two different names; Serverless is a genuinely different one, with its own compute plane, its own storage, and no VNet for you to get wrong. No native multi-region capability; Databricks' own managed disaster recovery is a distinct, opt-in feature, not something every workspace gets.
Single-zone deployment or non-AZ region
In regions without Availability Zone support, or when compute capacity is concentrated in one zone under high demand, the control plane and cluster nodes aren't spread across failure domains.
- Replication
- None
- RPO
- 0 for control plane metadata; compute-plane cached data is ephemeral and expected to be lost on node failure
- RTO
- Cluster/job restart time if the driver node is lost — no committed figure for this configuration
- Failover trigger
- Automatic
- Relative cost
- No extra cost
Baseline — you pay for the same VM count regardless of zone placement; this option just doesn't spread them.
Shows Unknown at every tier — there's no fixed RPO/RTO for this option; it depends entirely on an architecture you'd have to design and build.
Zone-redundant control plane + compute, ZRS/GZRS workspace storage
Control plane databases replicate synchronously across zones automatically for every workspace type. Below that, zone spread depends on which compute plane is doing the work. Classic compute (Classic/Hybrid workspaces) runs inside your own workspace VNet — Databricks still auto-distributes nodes across zones, but only if your VNet actually has subnets and capacity in each zone, which is on you to provision correctly. Serverless compute (in Serverless workspaces, or run alongside classic compute in a Hybrid workspace) has zone selection handled entirely by Databricks in its own account — no VNet to design or misconfigure. Storage redundancy follows the same split: Classic/Hybrid workspaces expose a customer-visible workspace storage account you can set to ZRS/GZRS instead of the GRS default; Serverless workspaces use Databricks-managed default storage instead, without that same direct SKU choice.
- Replication
- Synchronous
- RPO
- 0 — control plane expects no data loss during a zone outage; compute-plane cached data is ephemeral and recomputed
- RTO
- Control plane fails over to healthy zones within about 15 minutes. Losing a cluster's driver node restarts the whole cluster; worker-node loss alone replaces faster.
- Failover trigger
- Automatic
- Relative cost
- Low cost premium
Compute cost is unaffected by zone spread either way. On Classic/Hybrid workspaces, switching workspace storage from the GRS default to ZRS/GZRS changes storage pricing — a modest, storage-only premium.
The 15-minute figure is for the control plane. A lost driver node restarts the whole cluster and its running job — plan job checkpointing accordingly, not just zone redundancy. "Classic" and "Hybrid" are the same workspace type under different names (see gotchas) — the meaningful split for zone resiliency is Classic/Hybrid vs. Serverless compute, not Classic vs. Hybrid.
Managed disaster recovery (gated), or a customer-built secondary workspace
Databricks is a single-region service with no built-in multi-region capability. Managed disaster recovery is a distinct, gated feature — you apply for access through your Azure Databricks account team before you can use it. Once enabled, it continuously replicates Unity Catalog managed tables (Delta Lake, with data), external tables/volumes (metadata only, not the underlying data), views, functions, and permission grants to a secondary workspace, plus — optionally, off by default — workspace assets like notebooks, jobs, SQL warehouses, clusters, and their ACLs. It also offers an optional stable URL that always resolves to whichever workspace is currently primary. It requires the Premium plan, a separately-billed Mission Critical add-on, and serverless compute enabled on both workspaces (the secondary's serverless compute is what actually reads from the primary's storage during replication). Without it, or for what it doesn't cover, the standard pattern is a second workspace you build and sync yourself.
- Replication
- Asynchronous
- RPO
- No fixed number — Databricks doesn't publish a committed RPO even for managed DR. Replication is continuous, but data written after the last synced 'replication point' can be lost during an unplanned failover; you monitor actual RPO trends via a system table
- RTO
- Failover itself completes in minutes once triggered from the account console — but replicated clusters arrive terminated, SQL warehouses arrive stopped, and job schedules stay paused until you manually resume them
- Failover trigger
- Manual (customer-triggered)
- Relative cost
- High cost premium
A second workspace, its own compute, plus the Mission Critical add-on billed separately on both workspaces for as long as managed DR stays enabled — a real, ongoing cost on top of the duplicate infrastructure.
Managed DR narrows the DIY problem for what it covers, but the exclusion list is long — see gotchas — and you still decide when to fail over. "Minutes" describes the platform action, not the full return to a running, job-scheduled state.
Shows Unknown at every tier — there's no fixed RPO/RTO for this option; it depends entirely on an architecture you'd have to design and build.
Gotchas
Databricks has no multi-region capability out of the box
Every workspace is pinned to one region. Even with the managed disaster recovery feature, you're maintaining a second workspace and deciding when to fail over — it narrows the DIY problem, it doesn't remove it.
Managed disaster recovery is gated — you can't just turn it on
It isn't a self-service setting. Microsoft's own documentation states: "Managed DR is gated. Apply for access through your Azure Databricks account team." Any DR plan or SLA commitment that assumes managed DR is available the moment you decide you need it is wrong — budget approval and provisioning time into the plan, not just the failover procedure.
Managed DR's coverage has real gaps — ML, streaming, and secrets aren't replicated
The exclusion list is long: materialized views, streaming tables, Lakeflow pipelines, the actual data inside managed volumes (only their metadata replicates), Unity Catalog and workspace secrets, ML models, model serving endpoints, vector search indexes, Delta shares, published AI/BI dashboards (only drafts replicate), and Spark Structured Streaming jobs run outside a Lakeflow pipeline are all out of scope. Tables with row filters or column masks, and ABAC-tagged resources, are flagged as "failed to replicate" and block RPO until removed from scope. A workspace built around any of these still needs its own separately-designed DR plan for those specific pieces, even with managed DR enabled and working correctly for everything else.
Setup isn't instant: the Mission Critical add-on is a standing cost, and initial bootstrap can take two weeks
Managed DR requires the Premium plan plus the Mission Critical add-on enabled on both workspaces, billed separately at its own rate on all compute in each workspace for as long as it stays on. For a large workspace, the first replication cycle — the initial bootstrap that copies everything in scope to the secondary — can take up to two weeks before managed DR is even ready to use. Plan the timeline as carefully as the budget.
"Classic" and "Hybrid" are the same workspace type, not two different ones
Microsoft's own documentation states it plainly: "Classic workspaces are referred to as Hybrid workspaces in the Azure portal." There's no separate Classic offering to compare against Hybrid — they're identical, just named differently depending on whether you're reading Databricks docs (Classic) or looking at the Azure portal (Hybrid). The real, resiliency-relevant split is Classic/Hybrid vs. Serverless, which are genuinely different compute planes with different network and storage models.
Classic compute's zone spread depends on your VNet; serverless compute's doesn't
On a Classic/Hybrid workspace, classic compute resources run inside your own workspace VNet — Databricks only distributes cluster nodes across zones if that VNet actually has subnets and capacity in each zone, a customer-owned configuration step that can silently be wrong. Serverless compute (available in Serverless workspaces, or alongside classic compute in a Hybrid workspace) runs in a Databricks-managed compute plane with zone selection and replacement handled entirely by Databricks — there's no VNet to misconfigure in the first place.
A lost driver node restarts the whole cluster, not just a worker
Zone redundancy for compute means nodes get replaced automatically, but if the specific node lost is the driver, the entire cluster and its running job restart. Job checkpointing and idempotent design matter as much as the platform's zone spread.
Classic/Hybrid workspace storage defaults to GRS, not ZRS/GZRS
The workspace storage account Databricks auto-provisions on a Classic/Hybrid workspace starts on geo-redundant storage by default, not zone-redundant — you have to explicitly reconfigure it for zone protection. Serverless workspaces use Databricks-managed default storage instead, without that same customer-facing SKU choice — verify its current redundancy commitment directly with Databricks/Microsoft rather than assuming it matches what you'd pick for a classic workspace.