Exporter
Use these metrics to know whether the numbers you see are fresh, and to find which Proxmox VE call is slow or failing.
Scrape
Section titled “Scrape”Always exported.
| Metric | Type | Labels | Unit | Meaning |
|---|---|---|---|---|
cv4pve_scrape_duration_seconds |
gauge | - | seconds | Time the last scrape took to read the cluster |
cv4pve_scrape_last_success_timestamp_seconds |
gauge | - | Unix time | End of the last scrape in which every call succeeded |
cv4pve_scrape_errors_total |
counter | section |
- | Failed calls, by section: cluster (the cluster-wide calls), node (per-node calls), guest (balloon) |
cv4pve_scrape_duration_seconds is what to compare with the scrape_timeout of Prometheus. It grows
with the size of the cluster and the collectors that are on: see Settings
to bring it down.
A scrape with a failed call still answers with the other metrics, so up stays 1.
cv4pve_scrape_last_success_timestamp_seconds is what stops moving:
# No complete read of the cluster for 10 minutestime() - cv4pve_scrape_last_success_timestamp_seconds > 600
# Failed calls in the last 15 minutes, by sectionincrease(cv4pve_scrape_errors_total[15m]) > 0API instrumentation
Section titled “API instrumentation”Setting ApiInstrumentation (on in standard and full, off in fast).
| Metric | Type | Labels | Unit | Meaning |
|---|---|---|---|---|
cv4pve_api_request_duration_seconds |
histogram | method, endpoint |
seconds | Duration of each Proxmox VE API call |
cv4pve_api_request_errors_total |
counter | method, endpoint |
- | API calls answered with an HTTP error |
methodisGet, orCreatefor aPOST: only the balloon call, which uses the QEMU monitor.endpointis the API path with the variable parts replaced, so that each endpoint is one series:/nodes/{node}/status,/nodes/{node}/qemu/{vmid}/monitor,/cluster/resources.- The histogram buckets are 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10 and 30 seconds. As for any histogram,
Prometheus stores
_bucket,_sumand_countseries. - The login at the start of a scrape is not counted.
# Average duration of each endpoint over 5 minutesrate(cv4pve_api_request_duration_seconds_sum[5m]) / rate(cv4pve_api_request_duration_seconds_count[5m])
# 99th percentile per endpointhistogram_quantile(0.99, sum by (le, endpoint) (rate(cv4pve_api_request_duration_seconds_bucket[5m])))
# Endpoints that are failingincrease(cv4pve_api_request_errors_total[15m]) > 0On a large cluster this is the quickest way to see where a scrape spends its time, for example the
balloon (/nodes/{node}/qemu/{vmid}/monitor), which is called once per running VM.