Skip to content

Exporter

Use these metrics to know whether the numbers you see are fresh, and to find which Proxmox VE call is slow or failing.

Always exported.

Metric Type Labels Unit Meaning
cv4pve_scrape_duration_seconds gauge - seconds Time the last scrape took to read the cluster
cv4pve_scrape_last_success_timestamp_seconds gauge - Unix time End of the last scrape in which every call succeeded
cv4pve_scrape_errors_total counter section - Failed calls, by section: cluster (the cluster-wide calls), node (per-node calls), guest (balloon)

cv4pve_scrape_duration_seconds is what to compare with the scrape_timeout of Prometheus. It grows with the size of the cluster and the collectors that are on: see Settings to bring it down.

A scrape with a failed call still answers with the other metrics, so up stays 1. cv4pve_scrape_last_success_timestamp_seconds is what stops moving:

# No complete read of the cluster for 10 minutes
time() - cv4pve_scrape_last_success_timestamp_seconds > 600
# Failed calls in the last 15 minutes, by section
increase(cv4pve_scrape_errors_total[15m]) > 0

Setting ApiInstrumentation (on in standard and full, off in fast).

Metric Type Labels Unit Meaning
cv4pve_api_request_duration_seconds histogram method, endpoint seconds Duration of each Proxmox VE API call
cv4pve_api_request_errors_total counter method, endpoint - API calls answered with an HTTP error
  • method is Get, or Create for a POST: only the balloon call, which uses the QEMU monitor.
  • endpoint is the API path with the variable parts replaced, so that each endpoint is one series: /nodes/{node}/status, /nodes/{node}/qemu/{vmid}/monitor, /cluster/resources.
  • The histogram buckets are 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10 and 30 seconds. As for any histogram, Prometheus stores _bucket, _sum and _count series.
  • The login at the start of a scrape is not counted.
# Average duration of each endpoint over 5 minutes
rate(cv4pve_api_request_duration_seconds_sum[5m])
/ rate(cv4pve_api_request_duration_seconds_count[5m])
# 99th percentile per endpoint
histogram_quantile(0.99, sum by (le, endpoint) (rate(cv4pve_api_request_duration_seconds_bucket[5m])))
# Endpoints that are failing
increase(cv4pve_api_request_errors_total[15m]) > 0

On a large cluster this is the quickest way to see where a scrape spends its time, for example the balloon (/nodes/{node}/qemu/{vmid}/monitor), which is called once per running VM.