Alerts
An alert is a notification that VeloDB Cloud sends when a monitored metric meets a condition that you define. An alert rule connects a target, such as a warehouse, cluster, or job, to one or more metric conditions. For example, you can create a rule that alerts you when cluster CPU utilization remains above 80% for 15 minutes or when a Routine Load job accumulates error rows.
Alert rules help you detect conditions that need attention without continuously watching the Metrics page. The rule evaluates recent metric data at the interval you configure and triggers only when the condition remains true for the required duration. This evaluation period helps reduce notifications caused by short-lived metric changes. When the rule fires, VeloDB Cloud records an event in Alert History and sends the notification through the channels configured for the rule, such as email, in-site notification, SMS, webhook, PagerDuty, Slack, Lark, or DingTalk.
Use alerts to monitor different operational scopes:
- Warehouse alerts notify you about conditions that affect the warehouse as a whole, such as high CPU utilization, high memory usage, or unavailable nodes.
- Cluster alerts notify you about compute capacity and query performance for a specific cluster, such as cache hit rate, query latency, or compaction score.
- Job alerts notify you about a specific Routine Load, Warmup, MySQL, or PostgreSQL job, such as ingestion lag, pending work, or error rows.
You can start with VeloDB Cloud's recommended alert rules and then edit them, or create custom rules for the metrics and thresholds that matter to your workload.
VeloDB Cloud provides monitoring and alerting at no additional charge, except for SMS notifications.
View alert rules
To view alert rules:
- Log in to the VeloDB Cloud console.
- In the upper-left corner, select the warehouse that you want to monitor.
- In the left navigation pane, under Monitoring, click Alerts.
The Alert Rules tab lists existing rules with their rule name, warehouse or cluster, conditions, target, firing count for the last 7 days, and available actions. From the list, you can enable, disable, copy, edit, or delete a rule.
To review previous alert events, open the Alert History tab. You can filter the history by time, rule, or cluster.
Enable recommended alerts
Enabling recommended alerts lets you quickly set up basic alarm rules. These alarm rules apply to both current and future warehouses or clusters. If you don't know which metrics to monitor, use the recommended set of alert rules. You can edit the rules after you enable them.
The set includes the following rules. Warehouse-scoped rules apply to the entire warehouse. Cluster-scoped rules apply to each cluster.
Note:
You can see these recommended rules in the Alert Rules list only after you enable them.
| Scope | Metric | Condition | Run this rule every | Evaluate the last | Require the condition for |
|---|---|---|---|---|---|
| Warehouse | CPU Utilization | > 60% | 15 seconds | 30 seconds of data | 15 minutes |
| Warehouse | Memory Usage | > 80% | 15 seconds | 30 seconds of data | 15 minutes |
| Warehouse | Connection Number | > 150 | 15 seconds | 30 seconds of data | 15 minutes |
| Warehouse | Dead Node Count | > 0 | 15 seconds | 30 seconds of data | 1 minute |
| Warehouse | FE Process Dead Node Count | > 0 | 15 seconds | 30 seconds of data | 1 minute |
| Cluster | CPU Utilization | > 80% | 15 seconds | 30 seconds of data | 15 minutes |
| Cluster | Memory Usage | > 80% | 15 seconds | 30 seconds of data | 15 minutes |
| Cluster | Dead Node Count | > 0 | 15 seconds | 30 seconds of data | 1 minute |
| Cluster | Disk Cache Hit Rate | < 90% | 15 seconds | 30 seconds of data | 15 minutes |
| Cluster | Average Query Time | > 5,000 ms | 15 seconds | 300 seconds of data | 5 minutes |
| Cluster | P99 Query Latency | > 60,000 ms | 15 seconds | 300 seconds of data | 5 minutes |
| Cluster | Query Success Rate | < 90% | 15 seconds | 300 seconds of data | 5 minutes |
| Cluster | Compaction Score | > 1,500 | 15 seconds | 30 seconds of data | 15 minutes |
To enable the recommended set of alert rules:
- Log in to the VeloDB Cloud console.
- In the upper-left corner, select the warehouse that you want to monitor.
- In the left navigation pane, under Monitoring, click Alerts.
- In the upper-left corner, click Enable Recommended Alerts to create a recommended set of alert rules automatically.
You can then see all the recommended rules and edit them to change their conditions, duration, or notification channels.
Create an alert rule
Besides the recommended rules, you can create your own alert rules to monitor warehouse, cluster, or job metrics. To create an alert rule:
- Log in to the VeloDB Cloud console.
- In the upper-left corner, select the warehouse that you want to monitor.
- In the left navigation pane, under Monitoring, click Alerts.
- Click New Alert Rule, or copy an existing rule from the list to use it as a template.
An alert rule has four sections:
| Part | Description |
|---|---|
| Rule basics | Enter a rule name that is unique within the warehouse, then select the target. |
| Conditions | Define when the alert should trigger. Each condition compares a metric with a value by using an operator (>, >=, <, <=, ==, !=). Combine multiple conditions with and or or. |
| Evaluation | Control how often and how long the conditions are checked. The form evaluates a defined period of data and requires the condition to remain true for a defined duration before alerting. |
| Notification | Select one or more notification channels and configure how long to suppress repeat notifications for the same alert. For more information, see Notification channels. |
Rule basics
When you create a rule, you must select a target. The target determines which metrics are available for the rule.
A rule monitors one of the available target types:
- Compute monitors the selected warehouse or compute cluster. Available metrics depend on the selected target.
- Routine Load Job monitors a single Routine Load job. Go to the corresponding Routine Load Job details page to create alert rules.
- Warmup Job monitors a single Warmup job. Go to the corresponding Warmup Job details page to create alert rules.
- MySQL Job and PostgreSQL Job monitor the corresponding job types when they are available in your environment. Go to the corresponding MySQL Job and PostgreSQL Job details pages to create alert rules.
When you delete a target, its rules are retained but become inactive. Create compute rules from the Alerts page. Create job rules from the job's own page, not from the Alerts page (see Job metrics).
Conditions
You can define when this alert triggers with the available metrics for the selected target. Each condition compares a metric with a value by using an operator (>, >=, <, <=, ==, !=). Combine multiple conditions with and or or.
Evaluation
The evaluation section controls how often and how long the conditions are checked. The form evaluates a defined period of data and requires the condition to remain true for a defined duration before alerting.
Notification channels
Each alert rule can send notifications to one or more channels. Each channel delivers the alert message independently.
Email
Select the users who should receive notifications.
In-site Notification
Select the users who should receive notifications. VeloDB Cloud displays the alert in the console notification bell.
SMS
Select users or enter phone numbers directly. SMS notifications may incur additional charges.
Webhook
The Webhook channel forwards alert payloads to any external system that accepts HTTP callbacks. When an alert fires, VeloDB Cloud sends a POST request with a JSON body. Your endpoint must respond with HTTP 200. Otherwise, VeloDB Cloud retries the delivery.
PagerDuty
The PagerDuty channel routes alerts to PagerDuty incidents through a service integration key. To set up this channel:
- In PagerDuty, open Services, then Service Directory, and open the target service or create a new one.
- Open the Integrations tab, then click Add Integration.
- Search for Events API V2, select it, then click Create Service.
- Copy the Integration Key from the integration detail page.
- Paste the integration key into the alert channel configuration.
Slack Group
The Slack Group channel forwards alerts to Slack incoming webhook URLs. To use more than one URL, separate the URLs with a semicolon (;). To set up this channel:
- In Slack, open Apps, search for Incoming WebHooks, then click Add to Slack.
- Choose the channel that should receive alerts, then click Add Incoming WebHooks Integration.
- Copy the webhook URL into the alert channel configuration.
Lark Group
The Lark Group channel forwards alerts to a Lark or Feishu group bot. To set up this channel:
- In the target group, open Settings, choose BOTs, click Add Bot, then select Custom Bot.
- Give the bot a name and description, then click Next.
- Copy the webhook URL into the alert channel configuration.
Dingtalk Group
The Dingtalk Group channel forwards alerts to a DingTalk group robot. See DingTalk's guide for the full procedure. The main steps are:
- In the target DingTalk group, open Group Settings, then Group Assistant, then Add Robot, and choose Custom.
- Set a profile picture, name, and security settings. For security, use Custom Keywords and enter
alert. - Accept the terms, then click Finished.
- Copy the webhook URL into the alert channel configuration.
Note:
The Slack, Lark, and DingTalk channels let you restrict message sources with a webhook IP allowlist. The VeloDB Cloud server IP is
3.222.235.198.
Available metrics
You can create alerts for any warehouse or cluster metric shown on the Metrics page. See that page for each metric's meaning, unit, scope, and minimum version.
Job metrics are specific to a single Routine Load or Warmup job and are listed below.
Job metrics
Create job rules from the job's page, not from the Alerts page:
- Routine Load: In the left navigation pane, under INGESTION, click Import Data, click a job name to open its job page, then click Create Alert. For more information, see Routine Load.
- Warmup: In the left navigation pane, under COMPUTE, click Warmup Jobs, click a Job ID to open its job page, then click Create Alert. For more information, see Warmup.
Use job metrics to detect problems in an individual job, such as a Routine Load job that falls behind or accumulates errors.
| Metric | Job type | Unit |
|---|---|---|
| Routine Load Job Error Rows | Routine Load | rows |
| Routine Load Job Lag | Routine Load | rows |
| Routine Load Job Aborted Transactions | Routine Load | tasks |
| Warmup Job Trigger Gap | Warmup | ms |
| Warmup Job Pending Size | Warmup | bytes |
These job metrics are available from v4.1.8.
See also
- Metrics: review the metrics you can alert on.
- Metrics API: scrape metrics into your own monitoring stack.