Metrics
Use Metrics in the VeloDB Cloud console to investigate warehouse and cluster health over a selected time range. Each metric appears on a time-series card with minimum, average, maximum, and current values. For external collection and dashboards, use the Metrics API.
View metrics in the console
- Log in to the VeloDB Cloud console.
- In the upper-left corner, select the warehouse to inspect.
- In the left navigation pane, under OBSERVABILITY, click Metrics.
Explore metrics
- Select Warehouse to view warehouse-level metrics, or select a cluster tab to view metrics for that cluster.
- Select a time interval, then use the date range picker to specify the period to investigate.
- Select Resource Metrics or Service Metrics.
- For cluster resource metrics, use Nodes to filter metrics by node.
- Inspect the Min, Avg, Max, and Current values below each chart. Hover over the chart to compare values at a specific time.
- Select the star icon on a metric card to add it to Star, where you can view important metrics together.
To refresh charts automatically, select 5 seconds, 10 seconds, 30 seconds, 1 minute, 5 minutes, or 15 minutes instead of off. Turn off automatic refresh when you compare a fixed incident window.
Metrics are split into two categories:
- Resource Metrics for compute, memory, storage, cache, network, and node health
- Service Metrics for query and ingestion workload behavior
The following tables list the metrics in each console view. A metric can appear in both warehouse and cluster views because the console calculates it at both levels. The Available from column shows the earliest supported core version. Initial version indicates that the metric has been available since the initial release.
Warehouse resource metrics
Warehouse resource metrics aggregate physical resource and node-health signals across the warehouse.
| Metric | Unit | Available from | What it shows |
|---|---|---|---|
| CPU Utilization | % | Initial version | Tracks CPU utilization across all warehouse nodes to identify sustained load or resource bottlenecks. |
| Memory Usage | GB | Initial version | Shows memory usage across all warehouse nodes to detect memory pressure. |
| Memory Utilization | % | Initial version | Shows memory utilization across all warehouse nodes to detect sustained memory pressure. |
| I/O Utilization | % | Initial version | Shows disk I/O utilization across warehouse nodes to identify storage bottlenecks. |
| Network Inbound/Outbound Throughput | MB/s | Initial version | Shows inbound and outbound network throughput for the warehouse to identify bandwidth bottlenecks. |
| Unreachable Node Count | Not available | Initial version | Shows the number of host-level unreachable nodes in the warehouse, based on host availability checks. |
| Inactive FE Service Node Count | Not available | Initial version | Shows the number of warehouse FE nodes whose service process is not active, helping identify systemd service failures or process exits. |
Warehouse service metrics
Warehouse service metrics aggregate query, ingestion, refresh, and virtual-cluster activity across the warehouse.
| Metric | Unit | Available from | What it shows |
|---|---|---|---|
| Queries per Second | QPS | Initial version | Shows queries processed by the warehouse per second to help assess throughput and peak demand. |
| Query Success Rate | % | Initial version | Shows the percentage of successful queries out of total queries, helping identify increases in query failures. |
| Rows Loaded per Second | Row/s | Initial version | Shows rows written per second by load jobs to track ingestion throughput. |
| Load Throughput | MB/s | Initial version | Shows data written per second by load jobs to track ingestion throughput. |
| Active Connections | Not available | Initial version | Shows active connections for the warehouse across all compute clusters. |
| Finished Load Job Rate | Not available | Initial version | Shows the recent load job completion rate. Sudden spikes or drops can indicate changes in workload or ingestion pipelines. |
| Average Load Job Latency | ms | Initial version | Shows average execution latency for warehouse load jobs to help assess how quickly ingested data becomes queryable. |
| Stream Load Request Rate | Not available | Initial version | Shows the Stream Load request rate for the warehouse to help monitor real-time ingestion traffic. |
| Broker Load Job Rate | Not available | Initial version | Shows the Broker Load job rate for the warehouse to help monitor batch ingestion traffic. |
| Insert Into Job Rate | Not available | Initial version | Shows the Insert Into write rate for the warehouse to help monitor SQL-based ingestion traffic. |
| Routine Load Jobs by Status | Not available | Initial version | Shows the number of Routine Load jobs by status for the warehouse to help track continuous ingestion health. |
| Routine Load Lag | Not available | 4.0.5 | Shows Routine Load consumption lag from systems such as Kafka to assess data freshness for real-time ingestion. |
| Routine Load Aborted Transactions | Not available | 4.0.5 | Shows Routine Load transactions aborted in the last 5 minutes. Rapid growth usually points to ingest failures, upstream issues, or configuration changes. |
| Routine Load Jobs in ABNORMAL_PAUSED State | Not available | 4.0.5 | Shows Routine Load jobs in the ABNORMAL_PAUSED state. This usually requires reviewing the pause reason and resuming the job after remediation. |
| Virtual Cluster Active-Standby Switches | Not available | 4.1.7 | Shows virtual cluster active-standby switch events in a recent window, with a minimum 10-minute lookback. Alert when the count is greater than 0. |
| Data Size | GB | Initial version | Shows historical data size trends for the warehouse in GB. |
| MV Refresh Skipped Tasks | tasks | 4.1.8 | Shows materialized view refresh tasks skipped in the lookback window to detect scheduling delays or backlog. |
| MV Refresh Queue Length | tasks | 4.1.8 | Shows running and pending materialized view refresh tasks to track refresh queue backlog. |
| MV Refresh Success Rate | % | 4.1.8 | Shows successful materialized view refresh tasks as a percentage of successful and failed tasks in the lookback window. Windows with no refresh tasks are shown as 100%. |
| MV Refresh Duration | ms | 4.1.8 | Shows P75 and P95 materialized view refresh durations derived from summary quantiles. |
Cluster resource metrics
Cluster resource metrics show physical resource, cache, storage, network, and node-health signals for the selected cluster. Use Nodes to include or exclude individual nodes.
| Metric | Unit | Available from | What it shows |
|---|---|---|---|
| CPU Utilization | % | Initial version | Tracks CPU utilization across cluster nodes to help you choose low-traffic windows for scaling or other resource-intensive operations. |
| Memory Usage | GB | Initial version | Shows memory usage across cluster nodes. Sustained high usage may indicate a need to tune workloads or scale out. |
| Memory Utilization | % | Initial version | Shows memory utilization across cluster nodes. Sustained high usage may indicate a need to tune workloads or scale out. |
| I/O Utilization | % | Initial version | Shows disk I/O utilization across cluster nodes. Sustained high values may affect query performance and indicate a need to scale or optimize data access. |
| Disk Cache Utilization | % | 4.0.4 | Shows local cache space utilization for the cluster to assess cache capacity. |
| Disk Cache Hit Rate | % | Initial version | Shows the local cache hit rate for the cluster to evaluate how effectively hot data is served from cache. |
| Disk Cache Read/Write IOPS | IOPS | Initial version | Shows local cache read and write IOPS for the cluster to monitor cache-layer request load. |
| Disk Cache Read/Write Throughput | MB/s | Initial version | Shows local cache read and write throughput for the cluster to assess cache-layer data transfer load. |
| Remote Storage Read/Write IOPS | IOPS | 4.0.4 | Shows remote storage read and write IOPS for the cluster to monitor object storage request load. |
| Remote Storage Read/Write Throughput | MB/s | 4.0.4 | Shows remote storage read and write throughput for the cluster to assess object storage data transfer load. |
| Network Inbound/Outbound Throughput | MB/s | Initial version | Shows inbound and outbound network throughput for the cluster to identify bandwidth bottlenecks. |
| Unreachable Node Count | Not available | Initial version | Shows the number of host-level unreachable nodes in the cluster, based on host availability checks. |
Cluster service metrics
Cluster service metrics show query demand, query latency, ingestion throughput, and compaction pressure for the selected cluster.
| Metric | Unit | Available from | What it shows |
|---|---|---|---|
| Queries per Second | QPS | Initial version | Shows queries processed by the cluster per second to help assess throughput and peak demand. |
| Query Success Rate | % | Initial version | Shows the percentage of successful queries out of total queries, helping identify increases in query failures. |
| Average Query Latency | ms | Initial version | Shows average query latency for the cluster to identify cluster-wide latency increases. |
| P99 Query Latency | ms | Initial version | Shows P99 query latency for the cluster to identify tail latency and the risk of slow queries. |
| Rows Loaded per Second | Row/s | Initial version | Shows rows written per second by load jobs to track ingestion throughput. |
| Load Throughput | MB/s | Initial version | Shows data written per second by load jobs to track ingestion throughput. |
| Compaction Score | Not available | Initial version | Shows compaction pressure for the cluster. Higher scores indicate heavier compaction load that may affect write or query performance. |
Job-level metrics for individual Routine Load and Warmup jobs appear on the corresponding job page, not on Metrics. To create alerts for warehouse, cluster, or job metrics, see Alerts.
See also
- Metrics API: scrape these metrics into your own Prometheus or Grafana stack.
- Alerts: get notified when a metric crosses a threshold.