Skip to main content

Metrics

Use Metrics in the VeloDB Cloud console to investigate warehouse and cluster health over a selected time range. Each metric appears on a time-series card with minimum, average, maximum, and current values. For external collection and dashboards, use the Metrics API.

View metrics in the console​

  1. Log in to the VeloDB Cloud console.
  2. In the upper-left corner, select the warehouse to inspect.
  3. In the left navigation pane, under OBSERVABILITY, click Metrics.

Explore metrics​

  1. Select Warehouse to view warehouse-level metrics, or select a cluster tab to view metrics for that cluster.
  2. Select a time interval, then use the date range picker to specify the period to investigate.
  3. Select Resource Metrics or Service Metrics.
  4. For cluster resource metrics, use Nodes to filter metrics by node.
  5. Inspect the Min, Avg, Max, and Current values below each chart. Hover over the chart to compare values at a specific time.
  6. Select the star icon on a metric card to add it to Star, where you can view important metrics together.

To refresh charts automatically, select 5 seconds, 10 seconds, 30 seconds, 1 minute, 5 minutes, or 15 minutes instead of off. Turn off automatic refresh when you compare a fixed incident window.

Metrics are split into two categories:

  • Resource Metrics for compute, memory, storage, cache, network, and node health
  • Service Metrics for query and ingestion workload behavior

The following tables list the metrics in each console view. A metric can appear in both warehouse and cluster views because the console calculates it at both levels. The Available from column shows the earliest supported core version. Initial version indicates that the metric has been available since the initial release.

Warehouse resource metrics​

Warehouse resource metrics aggregate physical resource and node-health signals across the warehouse.

MetricUnitAvailable fromWhat it shows
CPU Utilization%Initial versionTracks CPU utilization across all warehouse nodes to identify sustained load or resource bottlenecks.
Memory UsageGBInitial versionShows memory usage across all warehouse nodes to detect memory pressure.
Memory Utilization%Initial versionShows memory utilization across all warehouse nodes to detect sustained memory pressure.
I/O Utilization%Initial versionShows disk I/O utilization across warehouse nodes to identify storage bottlenecks.
Network Inbound/Outbound ThroughputMB/sInitial versionShows inbound and outbound network throughput for the warehouse to identify bandwidth bottlenecks.
Unreachable Node CountNot availableInitial versionShows the number of host-level unreachable nodes in the warehouse, based on host availability checks.
Inactive FE Service Node CountNot availableInitial versionShows the number of warehouse FE nodes whose service process is not active, helping identify systemd service failures or process exits.

Warehouse service metrics​

Warehouse service metrics aggregate query, ingestion, refresh, and virtual-cluster activity across the warehouse.

MetricUnitAvailable fromWhat it shows
Queries per SecondQPSInitial versionShows queries processed by the warehouse per second to help assess throughput and peak demand.
Query Success Rate%Initial versionShows the percentage of successful queries out of total queries, helping identify increases in query failures.
Rows Loaded per SecondRow/sInitial versionShows rows written per second by load jobs to track ingestion throughput.
Load ThroughputMB/sInitial versionShows data written per second by load jobs to track ingestion throughput.
Active ConnectionsNot availableInitial versionShows active connections for the warehouse across all compute clusters.
Finished Load Job RateNot availableInitial versionShows the recent load job completion rate. Sudden spikes or drops can indicate changes in workload or ingestion pipelines.
Average Load Job LatencymsInitial versionShows average execution latency for warehouse load jobs to help assess how quickly ingested data becomes queryable.
Stream Load Request RateNot availableInitial versionShows the Stream Load request rate for the warehouse to help monitor real-time ingestion traffic.
Broker Load Job RateNot availableInitial versionShows the Broker Load job rate for the warehouse to help monitor batch ingestion traffic.
Insert Into Job RateNot availableInitial versionShows the Insert Into write rate for the warehouse to help monitor SQL-based ingestion traffic.
Routine Load Jobs by StatusNot availableInitial versionShows the number of Routine Load jobs by status for the warehouse to help track continuous ingestion health.
Routine Load LagNot available4.0.5Shows Routine Load consumption lag from systems such as Kafka to assess data freshness for real-time ingestion.
Routine Load Aborted TransactionsNot available4.0.5Shows Routine Load transactions aborted in the last 5 minutes. Rapid growth usually points to ingest failures, upstream issues, or configuration changes.
Routine Load Jobs in ABNORMAL_PAUSED StateNot available4.0.5Shows Routine Load jobs in the ABNORMAL_PAUSED state. This usually requires reviewing the pause reason and resuming the job after remediation.
Virtual Cluster Active-Standby SwitchesNot available4.1.7Shows virtual cluster active-standby switch events in a recent window, with a minimum 10-minute lookback. Alert when the count is greater than 0.
Data SizeGBInitial versionShows historical data size trends for the warehouse in GB.
MV Refresh Skipped Taskstasks4.1.8Shows materialized view refresh tasks skipped in the lookback window to detect scheduling delays or backlog.
MV Refresh Queue Lengthtasks4.1.8Shows running and pending materialized view refresh tasks to track refresh queue backlog.
MV Refresh Success Rate%4.1.8Shows successful materialized view refresh tasks as a percentage of successful and failed tasks in the lookback window. Windows with no refresh tasks are shown as 100%.
MV Refresh Durationms4.1.8Shows P75 and P95 materialized view refresh durations derived from summary quantiles.

Cluster resource metrics​

Cluster resource metrics show physical resource, cache, storage, network, and node-health signals for the selected cluster. Use Nodes to include or exclude individual nodes.

MetricUnitAvailable fromWhat it shows
CPU Utilization%Initial versionTracks CPU utilization across cluster nodes to help you choose low-traffic windows for scaling or other resource-intensive operations.
Memory UsageGBInitial versionShows memory usage across cluster nodes. Sustained high usage may indicate a need to tune workloads or scale out.
Memory Utilization%Initial versionShows memory utilization across cluster nodes. Sustained high usage may indicate a need to tune workloads or scale out.
I/O Utilization%Initial versionShows disk I/O utilization across cluster nodes. Sustained high values may affect query performance and indicate a need to scale or optimize data access.
Disk Cache Utilization%4.0.4Shows local cache space utilization for the cluster to assess cache capacity.
Disk Cache Hit Rate%Initial versionShows the local cache hit rate for the cluster to evaluate how effectively hot data is served from cache.
Disk Cache Read/Write IOPSIOPSInitial versionShows local cache read and write IOPS for the cluster to monitor cache-layer request load.
Disk Cache Read/Write ThroughputMB/sInitial versionShows local cache read and write throughput for the cluster to assess cache-layer data transfer load.
Remote Storage Read/Write IOPSIOPS4.0.4Shows remote storage read and write IOPS for the cluster to monitor object storage request load.
Remote Storage Read/Write ThroughputMB/s4.0.4Shows remote storage read and write throughput for the cluster to assess object storage data transfer load.
Network Inbound/Outbound ThroughputMB/sInitial versionShows inbound and outbound network throughput for the cluster to identify bandwidth bottlenecks.
Unreachable Node CountNot availableInitial versionShows the number of host-level unreachable nodes in the cluster, based on host availability checks.

Cluster service metrics​

Cluster service metrics show query demand, query latency, ingestion throughput, and compaction pressure for the selected cluster.

MetricUnitAvailable fromWhat it shows
Queries per SecondQPSInitial versionShows queries processed by the cluster per second to help assess throughput and peak demand.
Query Success Rate%Initial versionShows the percentage of successful queries out of total queries, helping identify increases in query failures.
Average Query LatencymsInitial versionShows average query latency for the cluster to identify cluster-wide latency increases.
P99 Query LatencymsInitial versionShows P99 query latency for the cluster to identify tail latency and the risk of slow queries.
Rows Loaded per SecondRow/sInitial versionShows rows written per second by load jobs to track ingestion throughput.
Load ThroughputMB/sInitial versionShows data written per second by load jobs to track ingestion throughput.
Compaction ScoreNot availableInitial versionShows compaction pressure for the cluster. Higher scores indicate heavier compaction load that may affect write or query performance.

Job-level metrics for individual Routine Load and Warmup jobs appear on the corresponding job page, not on Metrics. To create alerts for warehouse, cluster, or job metrics, see Alerts.

See also​

  • Metrics API: scrape these metrics into your own Prometheus or Grafana stack.
  • Alerts: get notified when a metric crosses a threshold.