Skip to main content

Disaster Recovery

Disaster recovery (DR) restores a workload when its primary region is unavailable. High Availability protects against failures within one region, and VeloDB Cloud backups are stored in the warehouse region. Neither is a substitute for a recoverable copy in another region.

VeloDB Cloud does not transparently replicate a warehouse and all its configuration across regions. Build cross-region DR by exporting required data to object storage in another region, or by having the application write to warehouses in both regions. You must also prepare the target warehouse, security configuration, network connectivity, and application cutover.

Plan the recovery scope​

Start with the business workload, not the warehouse. Identify everything required to serve that workload in the recovery region:

  • Data: databases, tables, partitions, and external data required by the workload.
  • SQL objects: table definitions, views, materialized views, catalogs, and other dependencies that are not recreated by a data-only export.
  • Access: warehouse SQL users and roles, credentials, cloud identities, and secrets required by applications and integrations.
  • Connectivity: public allowlists, private endpoints, DNS or service discovery, and firewall rules for the recovery region. Private endpoints and regional service URLs cannot be reused across regions.
  • Compute: cluster size, CPU architecture, cache capacity, and scaling configuration required to handle recovery traffic.
  • Application state: connection strings, retry behavior, scheduled jobs, ingestion checkpoints, and the procedure for pausing or redirecting writes.

Record the creation order for dependencies. For example, create the warehouse and network path before importing data, and create referenced databases and tables before restoring dependent views.

Choose a DR strategy​

StrategyRecovery point objectiveRecovery time objectiveUse whenMain tradeoffs
Scheduled export and importBased on the latest completed export that is available in the recovery region. Changes after that point must be replayed or are lost.Includes incident detection, recovery warehouse provisioning, configuration, data import, validation, and application cutover.Some data loss and a longer recovery time are acceptable.Lower standby cost, but recovery time grows with data volume and configuration complexity.
Pre-provisioned recovery warehouse with scheduled export and importUses the latest completed export, as above.Reduces provisioning and connectivity setup time. Data import and validation still determine much of the RTO.Infrastructure must be ready before an incident, but continuous dual-write is not required.Adds ongoing compute and administration cost. The recovery copy is not continuously current.
Application dual-writeDepends on the application's delivery guarantees, lag, retry logic, and reconciliation process.Can be shorter because the recovery warehouse already contains recent data and can be configured to serve read traffic.The workload requires a low RPO and RTO, and the application can coordinate writes to two independent warehouses.Highest cost and application complexity. Partial writes, ordering differences, and data divergence must be detected and reconciled.

Do not assume dual-write provides zero RPO. A request can succeed in one region and fail in the other. Define which region is authoritative, use stable event identifiers or another deduplication method, monitor replication lag, and provide a reconciliation process.

Export data to another region​

Export the databases and tables required by the recovery workload to object storage that remains accessible if the primary region is unavailable. Separately replicate or make available any external data required by the workload. The effective recovery point is the latest export that completed successfully and can be read from the recovery region.

Use Export Overview to choose between EXPORT and SELECT INTO OUTFILE. Select a format that preserves the data types and precision required by the target tables. Keep schema definitions and other SQL objects under version control because exported data files do not represent every warehouse-level dependency.

For each export cycle:

  1. Export the required data to a versioned or otherwise immutable destination path.
  2. Verify that the export task completed and that the expected files are present.
  3. Record the recovery point, object list, schema version, and any ingestion checkpoint needed to resume writes.
  4. Confirm that credentials and network policies in the recovery region can read the exported files.
  5. Retain enough completed exports to recover when the latest copy is incomplete or contains an application-level data error.

The export interval alone is not the RPO. A long-running or failed export increases the age of the latest usable recovery point. Monitor completion status and recovery-point age rather than only checking that the export was scheduled.

Prepare the recovery region​

Prepare regional resources before an incident when the target RTO does not allow time to build them during recovery:

  1. Confirm that VeloDB Cloud and the required warehouse features are available in the target region.
  2. Create or document how to create the recovery warehouse and size its clusters for the critical workload.
  3. Configure regional network connectivity and application access. See Connect to VeloDB Cloud.
  4. Recreate warehouse SQL users, roles, catalogs, and other required configuration without storing secrets in the runbook.
  5. Verify access to the cross-region object storage location and its encryption keys.
  6. Prepare an application endpoint or configuration that can be switched to the recovery warehouse.

Configuration can drift between regions even when the exported data is current. Review the recovery configuration after changes to schemas, permissions, integrations, networking, or warehouse sizing.

Recover after a regional outage​

Use a controlled failover process to avoid conflicting writes and an incomplete recovery:

  1. Declare the incident and determine the latest usable recovery point.
  2. If possible, stop writes to the primary region. For dual-write, decide which region is authoritative before accepting new writes.
  3. Provision or start the recovery warehouse, then apply its network, access, schema, and integration configuration.
  4. Import the exported data. See Load Overview for available import methods.
  5. Validate row counts, freshness, critical queries, permissions, and application dependencies.
  6. Redirect application traffic and ingestion to the recovery warehouse. Monitor errors, latency, and write progress.
  7. Record the actual recovery point and recovery time, including any data that must be replayed or reconciled.

Warning:

Do not direct writes to two regions during an unplanned recovery unless the application has a defined conflict-resolution process. Independent writes can cause duplicate, missing, or inconsistent data when you later consolidate the regions.

Test and maintain the DR plan​

Run recovery exercises with production-scale data in an isolated recovery warehouse. Plan for the additional compute and storage costs, and delete temporary resources when the exercise is complete. A successful export does not prove that the data can be imported within the target RTO or that the recovered application works.

During each exercise:

  • Restore into an isolated recovery warehouse without affecting production traffic.
  • Measure warehouse provisioning, configuration, import, validation, and application cutover time separately.
  • Verify data completeness at the intended recovery point and test critical read and write paths.
  • Confirm that the recovery region has enough capacity and that credentials, encryption keys, and private connectivity work.
  • Test failback as a separate procedure. Decide how writes made in the recovery region return to the original region before redirecting traffic.
  • Update the runbook after changes and after every incident or exercise.

See also​

  • Export Overview describes the export methods and commands.
  • Load Overview describes the available import methods.
  • High Availability explains how multi-zone deployment survives the loss of a single availability zone.
  • Reliability compares disaster recovery with high availability and backups.