Reliability
Reliability is the ability of a workload to continue serving users, preserve data, and recover within an agreed time after a failure. In VeloDB Cloud, reliability is a shared responsibility:
- VeloDB Cloud manages the service infrastructure, shared object storage, and the platform mechanisms for high availability, backup, and recovery.
- You configure the warehouse topology, backup and export schedules, application connection behavior, and recovery procedures that determine how your workload responds to a failure.
Use the following design sequence for each production workload:
- Define the business impact. Identify which queries and writes are critical, how long they can be unavailable, and how much data the business can recreate. Convert those requirements into a Recovery Time Objective (RTO) and Recovery Point Objective (RPO).
- Separate failure domains. A cluster failure, an availability zone failure, a bad write, and a regional outage require different protections. Do not use a same-region backup as the only protection against a regional outage.
- Remove single points of failure. Use two clusters in different availability zones as a primary-standby virtual cluster when the workload cannot wait for manual cluster replacement or applications need an automatically selected active endpoint.
- Make recovery repeatable. Document who declares an incident, how applications switch endpoints, which warehouse and data to restore, and how you verify the recovered result. A recovery option that has not been tested is an assumption, not an RTO.
- Detect and improve. Monitor warehouse and cluster health, query failures, backup and export task status, and recovery-point age. Review incidents and recovery drills to adjust capacity, schedules, and runbooks.
The following sections describe the VeloDB Cloud controls that implement this sequence. High availability reduces interruption within a region, backups and the Recycle Bin address data mistakes, and disaster recovery addresses regional loss. These controls complement one another and are not interchangeable.
RPO and RTO
Recovery Point Objective (RPO) is the maximum acceptable amount of data loss, expressed as the time between the latest recoverable point and the failure. For example, an RPO of one hour means that you may lose up to one hour of writes.
Recovery Time Objective (RTO) is the maximum acceptable time to restore service after a failure. It includes failure detection, recovery, and any data restore or import. Restore and import operations usually take longer as the data volume increases.
Operational readiness
Before a production launch or a high-risk change, confirm the following:
| Area | What to confirm |
|---|---|
| Workload | Critical queries, write paths, dependencies, and the owner responsible for recovery are documented. |
| Capacity | The standby or recovery warehouse can handle the expected workload, and the target region has the required service availability. |
| Data protection | Backup and cross-region export plans have completed successfully, and their latest completed recovery points meet the workload's RPO. |
| Connectivity | Applications can connect to the active or recovery warehouse, and retry or endpoint-switching behavior is tested. |
| Observability | Alerts and dashboards cover warehouse and cluster health, failed tasks, query errors, and recovery progress. Use Metrics, Alerts, and the Metrics API where you need external monitoring. |
| Exercise | A restore or failover drill has measured detection, provisioning, cache warmup, data restore, and application cutover time. |
Choose a protection
VeloDB Cloud provides different protections for different failure scopes. Use the following comparison to choose the protection that matches the failure you need to tolerate and the recovery point and recovery time you require.
| Failure scope | Protection | What it does | Recovery model and RPO/RTO considerations |
|---|---|---|---|
| One availability zone in a region | High Availability | Keeps warehouse metadata and shared data available across availability zones. | A warehouse with one cluster requires manual cluster replacement after a cluster failure. A primary-standby virtual cluster fails over to a synchronized standby automatically. Committed writes in shared object storage are preserved, but in-flight requests can fail. |
| Accidental drop, truncate, or other bad write | Recycle Bin or Backups | Recycle Bin restores recently removed SQL objects. Backups restore data from a retained recovery point. | Recycle Bin retention is resource-dependent and is best effort. Use a backup when you need a retained recovery point. The effective backup RPO is based on the most recent completed backup, and restore time depends on the data and warehouse resources. |
| Entire region unavailable | Disaster Recovery | Keeps a copy of data in another region and provides a recovery path there. | Cross-region export and import provide an RPO based on the export schedule and an RTO based on export, warehouse provisioning, and import time. Application dual-write can reduce both objectives, at the cost of operating two warehouses. |
High availability
Multi-AZ deployment protects against the loss of one availability zone within a region. Select a region that supports multi-AZ deployment when you create a warehouse. The warehouse uses shared object storage, but compute availability depends on how many clusters you deploy:
- Use one cluster when manual recovery is acceptable. Create a replacement cluster in another availability zone after a failure. The RTO depends on when you act and how long the cluster takes to start and warm its cache.
- Use two clusters in different availability zones as a primary-standby virtual cluster when you need automatic failover. The standby is synchronized in real time, but failover is not instantaneous because VeloDB Cloud must first detect the failure.
Recycle Bin
Run RECOVER to restore a database, table, or partition that you dropped or truncated while it remains in the recycle bin. VeloDB Cloud may remove recycle-bin entries to reclaim resources, so this is not a guaranteed recovery method. Configure Backups for recovery points that must be retained.
Backups
Backups copy data to object storage in the same region on a schedule or as a one-time task. You can restore from any retained backup. Choose the backup mode and schedule based on the granularity and recovery time you need. See Backups for the supported modes, version requirements, and restore procedure.
Disaster recovery
Multi-AZ deployment and backups remain within one region and do not protect against a regional outage. For cross-region protection, export the required databases to object storage in another region. After an outage, create a warehouse in that region, import the data, and redirect applications. See Disaster Recovery.
When importing after an outage would take too long, have the application write to warehouses in two regions. The second warehouse already contains a current copy and can continue serving traffic after a regional failure. This approach can reduce RPO and RTO, depending on replication and application cutover behavior, but it requires additional warehouse capacity and application-level write coordination.
Recommendations
- Set an RPO and RTO for each workload before choosing a protection. A single warehouse may need more than one protection because no option covers every failure scope.
- Schedule backups and cross-region exports so the interval between completed recovery points matches your RPO. Review task status and investigate failed or delayed runs.
- Run recovery drills with production-scale data. Measure detection, provisioning, cache warmup, export, import, and application cutover because these determine your actual RTO.