Skip to main content

Alerts

An alert is a notification that VeloDB Cloud sends when a monitored metric meets a condition that you define. An alert rule connects a target, such as a warehouse, cluster, or job, to one or more metric conditions. For example, you can create a rule that alerts you when cluster CPU utilization remains above 80% for 15 minutes or when a Routine Load job accumulates error rows.

Alert rules help you detect conditions that need attention without continuously watching the Metrics page. The rule evaluates recent metric data at the interval you configure and triggers only when the condition remains true for the required duration. This evaluation period helps reduce notifications caused by short-lived metric changes. When the rule fires, VeloDB Cloud records an event in Alert History and sends the notification through the channels configured for the rule, such as email, in-site notification, SMS, webhook, PagerDuty, Slack, Lark, or DingTalk.

Use alerts to monitor different operational scopes:

  • Warehouse alerts notify you about conditions that affect the warehouse as a whole, such as high CPU utilization, high memory usage, or unavailable nodes.
  • Cluster alerts notify you about compute capacity and query performance for a specific cluster, such as cache hit rate, query latency, or compaction score.
  • Job alerts notify you about a specific Routine Load, Warmup, MySQL, or PostgreSQL job, such as ingestion lag, pending work, or error rows.

You can start with VeloDB Cloud's recommended alert rules and then edit them, or create custom rules for the metrics and thresholds that matter to your workload.

VeloDB Cloud provides monitoring and alerting at no additional charge, except for SMS notifications.

View alert rules​

To view alert rules:

  1. Log in to the VeloDB Cloud console.
  2. In the upper-left corner, select the warehouse that you want to monitor.
  3. In the left navigation pane, under Monitoring, click Alerts.

The Alert Rules tab lists existing rules with their rule name, warehouse or cluster, conditions, target, firing count for the last 7 days, and available actions. From the list, you can enable, disable, copy, edit, or delete a rule.

To review previous alert events, open the Alert History tab. You can filter the history by time, rule, or cluster.

Enabling recommended alerts lets you quickly set up basic alarm rules. These alarm rules apply to both current and future warehouses or clusters. If you don't know which metrics to monitor, use the recommended set of alert rules. You can edit the rules after you enable them.

The set includes the following rules. Warehouse-scoped rules apply to the entire warehouse. Cluster-scoped rules apply to each cluster.

Note:

You can see these recommended rules in the Alert Rules list only after you enable them.

ScopeMetricConditionRun this rule everyEvaluate the lastRequire the condition for
WarehouseCPU Utilization> 60%15 seconds30 seconds of data15 minutes
WarehouseMemory Usage> 80%15 seconds30 seconds of data15 minutes
WarehouseConnection Number> 15015 seconds30 seconds of data15 minutes
WarehouseDead Node Count> 015 seconds30 seconds of data1 minute
WarehouseFE Process Dead Node Count> 015 seconds30 seconds of data1 minute
ClusterCPU Utilization> 80%15 seconds30 seconds of data15 minutes
ClusterMemory Usage> 80%15 seconds30 seconds of data15 minutes
ClusterDead Node Count> 015 seconds30 seconds of data1 minute
ClusterDisk Cache Hit Rate< 90%15 seconds30 seconds of data15 minutes
ClusterAverage Query Time> 5,000 ms15 seconds300 seconds of data5 minutes
ClusterP99 Query Latency> 60,000 ms15 seconds300 seconds of data5 minutes
ClusterQuery Success Rate< 90%15 seconds300 seconds of data5 minutes
ClusterCompaction Score> 1,50015 seconds30 seconds of data15 minutes

To enable the recommended set of alert rules:

  1. Log in to the VeloDB Cloud console.
  2. In the upper-left corner, select the warehouse that you want to monitor.
  3. In the left navigation pane, under Monitoring, click Alerts.
  4. In the upper-left corner, click Enable Recommended Alerts to create a recommended set of alert rules automatically.

You can then see all the recommended rules and edit them to change their conditions, duration, or notification channels.

Create an alert rule​

Besides the recommended rules, you can create your own alert rules to monitor warehouse, cluster, or job metrics. To create an alert rule:

  1. Log in to the VeloDB Cloud console.
  2. In the upper-left corner, select the warehouse that you want to monitor.
  3. In the left navigation pane, under Monitoring, click Alerts.
  4. Click New Alert Rule, or copy an existing rule from the list to use it as a template.

An alert rule has four sections:

PartDescription
Rule basicsEnter a rule name that is unique within the warehouse, then select the target.
ConditionsDefine when the alert should trigger. Each condition compares a metric with a value by using an operator (>, >=, <, <=, ==, !=). Combine multiple conditions with and or or.
EvaluationControl how often and how long the conditions are checked. The form evaluates a defined period of data and requires the condition to remain true for a defined duration before alerting.
NotificationSelect one or more notification channels and configure how long to suppress repeat notifications for the same alert. For more information, see Notification channels.

Rule basics​

When you create a rule, you must select a target. The target determines which metrics are available for the rule.

A rule monitors one of the available target types:

  • Compute monitors the selected warehouse or compute cluster. Available metrics depend on the selected target.
  • Routine Load Job monitors a single Routine Load job. Go to the corresponding Routine Load Job details page to create alert rules.
  • Warmup Job monitors a single Warmup job. Go to the corresponding Warmup Job details page to create alert rules.
  • MySQL Job and PostgreSQL Job monitor the corresponding job types when they are available in your environment. Go to the corresponding MySQL Job and PostgreSQL Job details pages to create alert rules.

When you delete a target, its rules are retained but become inactive. Create compute rules from the Alerts page. Create job rules from the job's own page, not from the Alerts page (see Job metrics).

Conditions​

You can define when this alert triggers with the available metrics for the selected target. Each condition compares a metric with a value by using an operator (>, >=, <, <=, ==, !=). Combine multiple conditions with and or or.

Evaluation​

The evaluation section controls how often and how long the conditions are checked. The form evaluates a defined period of data and requires the condition to remain true for a defined duration before alerting.

Notification channels​

Each alert rule can send notifications to one or more channels. Each channel delivers the alert message independently.

Email​

Select the users who should receive notifications.

In-site Notification​

Select the users who should receive notifications. VeloDB Cloud displays the alert in the console notification bell.

SMS​

Select users or enter phone numbers directly. SMS notifications may incur additional charges.

Webhook​

The Webhook channel forwards alert payloads to any external system that accepts HTTP callbacks. When an alert fires, VeloDB Cloud sends a POST request with a JSON body. Your endpoint must respond with HTTP 200. Otherwise, VeloDB Cloud retries the delivery.

PagerDuty​

The PagerDuty channel routes alerts to PagerDuty incidents through a service integration key. To set up this channel:

  1. In PagerDuty, open Services, then Service Directory, and open the target service or create a new one.
  2. Open the Integrations tab, then click Add Integration.
  3. Search for Events API V2, select it, then click Create Service.
  4. Copy the Integration Key from the integration detail page.
  5. Paste the integration key into the alert channel configuration.

Slack Group​

The Slack Group channel forwards alerts to Slack incoming webhook URLs. To use more than one URL, separate the URLs with a semicolon (;). To set up this channel:

  1. In Slack, open Apps, search for Incoming WebHooks, then click Add to Slack.
  2. Choose the channel that should receive alerts, then click Add Incoming WebHooks Integration.
  3. Copy the webhook URL into the alert channel configuration.

Lark Group​

The Lark Group channel forwards alerts to a Lark or Feishu group bot. To set up this channel:

  1. In the target group, open Settings, choose BOTs, click Add Bot, then select Custom Bot.
  2. Give the bot a name and description, then click Next.
  3. Copy the webhook URL into the alert channel configuration.

Dingtalk Group​

The Dingtalk Group channel forwards alerts to a DingTalk group robot. See DingTalk's guide for the full procedure. The main steps are:

  1. In the target DingTalk group, open Group Settings, then Group Assistant, then Add Robot, and choose Custom.
  2. Set a profile picture, name, and security settings. For security, use Custom Keywords and enter alert.
  3. Accept the terms, then click Finished.
  4. Copy the webhook URL into the alert channel configuration.

Note:

The Slack, Lark, and DingTalk channels let you restrict message sources with a webhook IP allowlist. The VeloDB Cloud server IP is 3.222.235.198.

Available metrics​

You can create alerts for any warehouse or cluster metric shown on the Metrics page. See that page for each metric's meaning, unit, scope, and minimum version.

Job metrics are specific to a single Routine Load or Warmup job and are listed below.

Job metrics​

Create job rules from the job's page, not from the Alerts page:

  • Routine Load: In the left navigation pane, under INGESTION, click Import Data, click a job name to open its job page, then click Create Alert. For more information, see Routine Load.
  • Warmup: In the left navigation pane, under COMPUTE, click Warmup Jobs, click a Job ID to open its job page, then click Create Alert. For more information, see Warmup.

Use job metrics to detect problems in an individual job, such as a Routine Load job that falls behind or accumulates errors.

MetricJob typeUnit
Routine Load Job Error RowsRoutine Loadrows
Routine Load Job LagRoutine Loadrows
Routine Load Job Aborted TransactionsRoutine Loadtasks
Warmup Job Trigger GapWarmupms
Warmup Job Pending SizeWarmupbytes

These job metrics are available from v4.1.8.

See also​

  • Metrics: review the metrics you can alert on.
  • Metrics API: scrape metrics into your own monitoring stack.