Skip to content

Monitoring and Observability

Every QHx control-plane component exposes a /metrics endpoint on a dedicated port. All endpoints use the prometheus named port and carry the label qhx.dev/prometheus-scrape: "true" on their Kubernetes Services.

ComponentDeployed PortFlag / Config
SPIRE server9998SPIRE telemetry block
SPIRE agentconfigurableSPIRE telemetry block
Manager9090--metrics-bind-address
Agent9091--metrics-addr
Proxy9092metricsAddr (config)

The Manager endpoint is plain HTTP (--metrics-secure=false in the cluster deployment). Set any address flag to 0 or "" to disable that component’s endpoint.

Apply the existing ServiceMonitor — it selects on qhx.dev/prometheus-scrape: "true" with a 30 s scrape interval and covers SPIRE instances, Manager, and Agent:

Terminal window
kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yaml

The Proxy runs as an injected sidecar, so its /metrics endpoint appears on user pods rather than a dedicated Service. Use a PodMonitor to scrape it:

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: qhx-proxy
namespace: monitoring
spec:
namespaceSelector:
any: true
selector:
matchLabels:
qhx.dev/proxy-injected: "true"
podMetricsEndpoints:
- port: qhx-proxy-metrics
path: /metrics
interval: 30s

Helm values:

manager:
metricsAddr: ":9090" # --metrics-bind-address
metricsSecure: false # --metrics-secure (set false for plain HTTP)
agent:
metricsAddr: ":9091"
proxy:
metricsAddr: ":9092"

Reconciliation

MetricTypeLabelsDescription
qhx_manager_reconcile_totalCountercontroller, resultReconcile attempts. controller: spire-instance | qhx-cluster-policy | bundle-exchange | license-status. result: success | error | requeue
qhx_manager_reconcile_duration_secondsHistogramcontrollerWall time of each reconcile call
qhx_manager_reconcile_requeue_after_secondsHistogramcontrollerRequested requeue delay when result is requeue

SPIRE Instance State

MetricTypeLabelsDescription
qhx_manager_spire_instances_totalGauge—Total desired SPIRE instances after policy evaluation
qhx_manager_spire_instances_readyGauge—SPIRE instances with at least one ready agent pod on the manager node
qhx_manager_spire_instances_deleted_totalCounter—Instances deleted after cooldown expiry or explicit removal
qhx_manager_spire_instances_pvc_rebuild_totalCounter—StatefulSet delete+recreate cycles triggered by immutable VolumeClaimTemplate changes
qhx_manager_spire_instance_hash_collisions_totalCounter—Deduplication events: two policies resolving to the same descriptor hash
qhx_manager_agent_sources_sync_errors_totalCounter—Failures syncing agent sources from the SPIFFE Workload API socket (non-fatal — reconcile proceeds)

Policy, Webhook, and Bundle Exchange

MetricTypeLabelsDescription
qhx_manager_policies_totalGaugekindPolicy object count. kind: cluster | namespace
qhx_manager_webhook_requests_totalCounteroperation, resource, resultAdmission webhook calls. result: allowed | denied | error
qhx_manager_webhook_duration_secondsHistogramoperation, resourceAdmission webhook latency
qhx_manager_bundle_exchange_totalCounterdirection, resultBundle exchange operations. direction: push | pull. result: success | error

License and Config

MetricTypeLabelsDescription
qhx_manager_license_validGauge—1 if at least one valid license is present, 0 otherwise
qhx_manager_config_reloads_totalCounterresult, sourceConfigMap/Secret reload events. result: success | error. source: configmap | secret (identifies what failed) | all (on success)
qhx_manager_config_watch_errors_totalCountersourceWatch loop failures (auto-restart after 5 s). source: configmap | secret
MetricTypeLabelsDescription
qhx_agent_workload_api_connectedGauge—1 when the SPIFFE Workload API socket is reachable, 0 after shutdown
qhx_agent_svid_renewals_totalCounterresultSVID renewal attempts. result: success | error
qhx_agent_svid_expiry_secondsGaugespiffe_idSeconds until the current SVID for the given SPIFFE ID expires
qhx_agent_workload_api_watch_errors_totalCounter—Errors from the Workload API watch stream

Connections and Traffic

MetricTypeLabelsDescription
qhx_proxy_connections_totalCounterprotocol, direction, resultConnection attempts. protocol: http | tcp | mqtt. direction: inbound | outbound. result: success | auth_failed | error
qhx_proxy_active_connectionsGaugeprotocol, directionCurrently open connections
qhx_proxy_bytes_received_totalCounterprotocolBytes received from clients
qhx_proxy_bytes_sent_totalCounterprotocolBytes forwarded to backends

Latency

MetricTypeLabelsDescription
qhx_proxy_tls_handshake_duration_secondsHistogramprotocol, rolemTLS handshake time including SPIFFE SVID validation. role: server | client. Currently only emitted for mqtt; HTTP and TCP delegate TLS to the standard library
qhx_proxy_request_duration_secondsHistogramprotocol, http_methodEnd-to-end request latency (HTTP only)

Authentication and Notary

MetricTypeLabelsDescription
qhx_proxy_auth_decisions_totalCounterresult, reason, protocol, roleSPIFFE identity auth outcomes. result: allowed | denied. reason: svid_invalid | policy_deny | ok. protocol: http | tcp | mqtt. role: server | client
qhx_proxy_notary_records_totalCounterresult, protocolAudit records written by the notary middleware. result: success | error. protocol: http | mqtt
qhx_proxy_notary_level_totalCounterlevelNotarization level distribution. level: workload | logRequest | signRequest
qhx_proxy_notary_db_errors_totalCounterop, entitybbolt database errors. op: put | get | init | close. entity: certificate_blob | workload_statement | receipt | db

MQTT Buffer

These metrics are only populated when the MQTT protocol is in use.

MetricTypeLabelsDescription
qhx_proxy_mqtt_buffer_messages_queuedGauge—Messages currently held in the store-and-forward buffer (not yet populated — requires middleware-layer instrumentation)
qhx_proxy_mqtt_buffer_messages_stored_totalCounter—Messages successfully forwarded to the upstream publisher
qhx_proxy_mqtt_buffer_messages_acked_totalCountersourceMessages for which a PUBACK was sent to the client. source: middleware (buffer layer handled the publish) | upstream (forwarded to broker and confirmed)
qhx_proxy_mqtt_buffer_replay_failures_totalCounter—Times the replay loop exhausted its backoff budget (not yet populated — requires middleware-layer instrumentation)
qhx_proxy_mqtt_upstream_connack_rejected_totalCounterreasonUpstream CONNACK rejections by reason code
qhx_proxy_mqtt_upstream_reconnections_totalCounterresultUpstream reconnection attempts. result: success | error

Sample alert rules:

groups:
- name: qhx-critical
rules:
- alert: QHxLicenseInvalid
expr: qhx_manager_license_valid == 0
for: 5m
labels:
severity: critical
annotations:
summary: "QHx license is not valid"
- alert: QHxManagerReconcileErrors
expr: rate(qhx_manager_reconcile_total{result="error"}[5m]) > 0.1
for: 2m
labels:
severity: warning
annotations:
summary: "Manager reconcile error rate elevated for {{ $labels.controller }}"
- alert: QHxAgentWorkloadAPIDown
expr: qhx_agent_workload_api_connected == 0
for: 1m
labels:
severity: critical
annotations:
summary: "QHx Agent cannot reach the SPIFFE Workload API"
- alert: QHxProxyAuthFailures
expr: rate(qhx_proxy_connections_total{result="auth_failed"}[5m]) > 0.5
for: 2m
labels:
severity: warning
annotations:
summary: "Elevated mTLS authentication failures on {{ $labels.protocol }} listener"

The Manager metrics endpoint defaults to HTTPS with authentication (--metrics-secure=true). For in-cluster Prometheus scraping without client-cert setup, set --metrics-secure=false in dev clusters.

Restrict Prometheus scrape access with a NetworkPolicy that only allows ingress from the monitoring namespace:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: qhx-metrics-scrape
namespace: qhx-system
spec:
podSelector:
matchLabels:
app.kubernetes.io/part-of: qhx
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: monitoring
ports:
- protocol: TCP
port: 9090
- protocol: TCP
port: 9091
- protocol: TCP
port: 9092
  1. Deploy ServiceMonitor:

    Terminal window
    kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yaml
  2. Verify metrics collection:

    Terminal window
    kubectl port-forward -n monitoring svc/prometheus 9090:9090
    # Open http://localhost:9090/targets and look for qhx-* entries
  3. Scrape a component directly:

    Terminal window
    # Manager (HTTP — set --metrics-secure=false in dev)
    kubectl port-forward -n qhx-system deploy/manager 9090:9090
    curl http://localhost:9090/metrics | grep qhx_manager
    # Agent
    kubectl port-forward -n qhx-system daemonset/qhx-agent 9091:9091
    curl http://localhost:9091/metrics | grep qhx_agent

Confirm metrics endpoint is accessible:

Terminal window
kubectl port-forward -n qhx-system deploy/manager 9090:9090
curl http://localhost:9090/metrics