HA
Use these metrics to know whether the HA manager is doing its job: that it has quorum, that every node
is online for HA, and that no HA resource is stuck in error, fence or recovery. None of this is in
the Proxmox VE metric server.
From /cluster/ha/resources and /cluster/ha/status/manager_status ·
setting Cluster.Ha (on in every profile; cached for 30 seconds in full).
| Metric | Type | Labels | Meaning |
|---|---|---|---|
cv4pve_ha_state |
gauge | sid, type, group, state |
1 for the state the HA manager reports for the resource now, 0 for the others |
cv4pve_ha_node_state |
gauge | node, state |
1 for the HA state of the node now, 0 for the others |
cv4pve_ha_quorate |
gauge | - | 1 if the HA manager has quorum |
A cluster with no HA resources has no cv4pve_ha_state series.
Resources
Section titled “Resources”Labels of cv4pve_ha_state:
| Label | Value |
|---|---|
sid |
The HA resource ID, e.g. vm:100 or ct:101 |
type |
vm or ct |
group |
The HA group of the resource, empty when it has none. Since Proxmox VE 9.0 groups are migrated to node affinity rules, and the label stays empty |
state |
One series per state the HA manager (CRM) can give a resource |
The states are those of the HA manager (the service state in manager_status), described in
the HA chapter of
the Proxmox VE documentation:
state |
Meaning |
|---|---|
started |
Active: the node starts it if it is not running, and restarts it if it fails |
stopped |
Stopped, confirmed by the node. A disabled resource is stopped too |
request_start, request_start_balance |
Start requested, not yet confirmed by the node |
request_stop |
Stop requested, the CRM waits for the node to confirm |
migrate |
Live migration to another node |
relocate |
Moving to another node with a stop and a start |
freeze |
Left untouched while its node reboots or the HA service restarts |
fence |
Its node is not in the quorate partition: waiting for the node to be fenced |
recovery |
Its node was fenced: waiting for a new node to run on |
error |
Disabled because of errors on the node: needs manual intervention |
The guest ID is inside sid: extract it into a vmid label to join with the guest metrics:
# HA resources in error, with the guest namelabel_replace(cv4pve_ha_state{state="error"} == 1, "vmid", "$1", "sid", "(?:vm|ct):(\\d+)") * on (vmid) group_left (name, node) cv4pve_guest_infocv4pve_ha_node_state has five series per node, one per state:
state |
Meaning |
|---|---|
online |
Online and member of the quorate partition |
maintenance |
In the quorate partition, but not able to do work (maintenance mode) |
unknown |
Not in the quorate partition, but possibly still running |
fence |
Must be fenced before its resources can be recovered |
gone |
No longer in the cluster member list, e.g. removed |