Skip to content

Monitoring and Observability

QHx provides comprehensive telemetry for security operations centers, enabling integration with enterprise monitoring systems and DoD cybersecurity operations infrastructure. All QHx components expose structured metrics and logs that integrate seamlessly with standard observability tools.

Target audience: Platform operators, security operations teams, system administrators.

QHx components expose Prometheus-compatible metrics for:

  • Authentication events - Node and workload attestation success/failure rates
  • Authorization events - Policy enforcement decisions and denials
  • Cryptographic operations - Certificate issuance, rotation, and lifecycle events
  • Component health - Service availability, resource utilization, and performance
  • Operational metrics - Request rates, latency, and throughput

Supported metric collectors:

  • Prometheus (recommended)
  • StatsD / DogStatsD
  • M3
  • OpenTelemetry Collector

QHx components emit structured JSON logs to stdout/stderr, enabling integration with standard Kubernetes logging stacks:

  • Security events - Authentication failures, authorization denials, policy violations
  • Operational events - Component lifecycle, configuration changes, errors
  • Audit events - Administrative actions, policy modifications, access attempts

Log fields include:

  • Timestamp (RFC3339)
  • Severity level
  • Component identifier
  • Event type and category
  • Contextual metadata (namespace, user, SPIFFE ID)

Supported log aggregators:

  • Fluent Bit / Fluentd (recommended)
  • Logstash
  • OpenTelemetry Collector
  • Splunk
  • Datadog

QHx supports integration with DoD cybersecurity operations infrastructure:

CSRMC (Cybersecurity Service Resource Management Center)

Section titled “CSRMC (Cybersecurity Service Resource Management Center)”

QHx telemetry is designed for CSRMC integration, providing:

  • Security event correlation
  • Incident detection and alerting
  • Compliance monitoring
  • Situational awareness

ACAS (Assured Compliance Assessment Solution)

Section titled “ACAS (Assured Compliance Assessment Solution)”

QHx supports ACAS vulnerability scanning:

  • Container image scanning
  • Configuration assessment
  • TLS/SSL certificate validation

QHx telemetry can feed HBSS for:

  • File integrity monitoring
  • Process monitoring
  • Network connection monitoring

Every QHx control-plane component exposes a /metrics endpoint on a dedicated port. All endpoints use the prometheus named port and carry the label qhx.dev/prometheus-scrape: "true" on their Kubernetes Services.

ComponentDeployed PortFlag / Config
SPIRE server9998SPIRE telemetry block
SPIRE agentconfigurableSPIRE telemetry block
Manager9090--metrics-bind-address
Agent9091--metrics-addr
Proxy9092metricsAddr (config)

The Manager endpoint is plain HTTP (--metrics-secure=false in the cluster deployment). Set any address flag to 0 or "" to disable that component’s endpoint.

Apply the existing ServiceMonitor — it selects on qhx.dev/prometheus-scrape: "true" with a 30 s scrape interval and covers SPIRE instances, Manager, and Agent:

Terminal window
kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yaml

The Proxy runs as an injected sidecar, so its /metrics endpoint appears on user pods rather than a dedicated Service. Use a PodMonitor to scrape it:

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: qhx-proxy
namespace: monitoring
spec:
namespaceSelector:
any: true
selector:
matchLabels:
qhx.dev/proxy-injected: "true"
podMetricsEndpoints:
- port: qhx-proxy-metrics
path: /metrics
interval: 30s

Helm values:

manager:
metricsAddr: ":9090" # --metrics-bind-address
metricsSecure: false # --metrics-secure (set false for plain HTTP)
agent:
metricsAddr: ":9091"
proxy:
metricsAddr: ":9092"

StatsD configuration:

# PKI Server configuration
telemetry {
Prometheus {
enabled = false
}
Statsd {
enabled = true
address = "statsd-agent.monitoring:8125"
}
}

Reconciliation

MetricTypeLabelsDescription
qhx_manager_reconcile_totalCountercontroller, resultReconcile attempts. controller: spire-instance | qhx-cluster-policy | bundle-exchange | license-status. result: success | error | requeue
qhx_manager_reconcile_duration_secondsHistogramcontrollerWall time of each reconcile call
qhx_manager_reconcile_requeue_after_secondsHistogramcontrollerRequested requeue delay when result is requeue

SPIRE Instance State

MetricTypeLabelsDescription
qhx_manager_spire_instances_totalGauge—Total desired SPIRE instances after policy evaluation
qhx_manager_spire_instances_readyGauge—SPIRE instances with at least one ready agent pod on the manager node
qhx_manager_spire_instances_deleted_totalCounter—Instances deleted after cooldown expiry or explicit removal
qhx_manager_spire_instances_pvc_rebuild_totalCounter—StatefulSet delete+recreate cycles triggered by immutable VolumeClaimTemplate changes
qhx_manager_spire_instance_hash_collisions_totalCounter—Deduplication events: two policies resolving to the same descriptor hash
qhx_manager_agent_sources_sync_errors_totalCounter—Failures syncing agent sources from the SPIFFE Workload API socket (non-fatal — reconcile proceeds)

Policy, Webhook, and Bundle Exchange

MetricTypeLabelsDescription
qhx_manager_policies_totalGaugekindPolicy object count. kind: cluster | namespace
qhx_manager_webhook_requests_totalCounteroperation, resource, resultAdmission webhook calls. result: allowed | denied | error
qhx_manager_webhook_duration_secondsHistogramoperation, resourceAdmission webhook latency
qhx_manager_bundle_exchange_totalCounterdirection, resultBundle exchange operations. direction: push | pull. result: success | error

License and Config

MetricTypeLabelsDescription
qhx_manager_license_validGauge—1 if at least one valid license is present, 0 otherwise
qhx_manager_config_reloads_totalCounterresult, sourceConfigMap/Secret reload events. result: success | error. source: configmap | secret (identifies what failed) | all (on success)
qhx_manager_config_watch_errors_totalCountersourceWatch loop failures (auto-restart after 5 s). source: configmap | secret
MetricTypeLabelsDescription
qhx_agent_workload_api_connectedGauge—1 when the SPIFFE Workload API socket is reachable, 0 after shutdown
qhx_agent_svid_renewals_totalCounterresultSVID renewal attempts. result: success | error
qhx_agent_svid_expiry_secondsGaugespiffe_idSeconds until the current SVID for the given SPIFFE ID expires
qhx_agent_workload_api_watch_errors_totalCounter—Errors from the Workload API watch stream

Connections and Traffic

MetricTypeLabelsDescription
qhx_proxy_connections_totalCounterprotocol, direction, resultConnection attempts. protocol: http | tcp | mqtt. direction: inbound | outbound. result: success | auth_failed | error
qhx_proxy_active_connectionsGaugeprotocol, directionCurrently open connections
qhx_proxy_bytes_received_totalCounterprotocolBytes received from clients
qhx_proxy_bytes_sent_totalCounterprotocolBytes forwarded to backends

Latency

MetricTypeLabelsDescription
qhx_proxy_tls_handshake_duration_secondsHistogramprotocol, rolemTLS handshake time including SPIFFE SVID validation. role: server | client. Currently only emitted for mqtt; HTTP and TCP delegate TLS to the standard library
qhx_proxy_request_duration_secondsHistogramprotocol, http_methodEnd-to-end request latency (HTTP only)

Authentication and Notary

MetricTypeLabelsDescription
qhx_proxy_auth_decisions_totalCounterresult, reason, protocol, roleSPIFFE identity auth outcomes. result: allowed | denied. reason: svid_invalid | policy_deny | ok. protocol: http | tcp | mqtt. role: server | client
qhx_proxy_notary_records_totalCounterresult, protocolAudit records written by the notary middleware. result: success | error. protocol: http | mqtt
qhx_proxy_notary_level_totalCounterlevelNotarization level distribution. level: workload | logRequest | signRequest
qhx_proxy_notary_db_errors_totalCounterop, entitybbolt database errors. op: put | get | init | close. entity: certificate_blob | workload_statement | receipt | db

MQTT Buffer

These metrics are only populated when the MQTT protocol is in use.

MetricTypeLabelsDescription
qhx_proxy_mqtt_buffer_messages_queuedGauge—Messages currently held in the store-and-forward buffer (not yet populated — requires middleware-layer instrumentation)
qhx_proxy_mqtt_buffer_messages_stored_totalCounter—Messages successfully forwarded to the upstream publisher
qhx_proxy_mqtt_buffer_messages_acked_totalCountersourceMessages for which a PUBACK was sent to the client. source: middleware (buffer layer handled the publish) | upstream (forwarded to broker and confirmed)
qhx_proxy_mqtt_buffer_replay_failures_totalCounter—Times the replay loop exhausted its backoff budget (not yet populated — requires middleware-layer instrumentation)
qhx_proxy_mqtt_upstream_connack_rejected_totalCounterreasonUpstream CONNACK rejections by reason code
qhx_proxy_mqtt_upstream_reconnections_totalCounterresultUpstream reconnection attempts. result: success | error

ConfigMap example:

apiVersion: v1
kind: ConfigMap
metadata:
name: fluent-bit-config
namespace: logging
data:
fluent-bit.conf: |
[INPUT]
Name tail
Path /var/log/containers/qhx-*_qhx-system_*.log
Parser docker
Tag kube.qhx.*
Refresh_Interval 5
Mem_Buf_Limit 5MB
Skip_Long_Lines On
[FILTER]
Name parser
Match kube.qhx.*
Key_Name log
Parser json
Reserve_Data On
[FILTER]
Name modify
Match kube.qhx.*
Add cluster_name ${CLUSTER_NAME}
Add environment ${ENVIRONMENT}
[OUTPUT]
Name es
Match kube.qhx.*
Host elasticsearch.monitoring
Port 9200
Index qhx-logs
Logstash_Format On
Logstash_Prefix qhx

QHx telemetry supports integration with alerting systems:

  • Prometheus AlertManager
  • Grafana
  • PagerDuty
  • Splunk
  • ServiceNow

Sample alert rules:

groups:
- name: qhx-critical
rules:
- alert: QHxLicenseInvalid
expr: qhx_manager_license_valid == 0
for: 5m
labels:
severity: critical
annotations:
summary: "QHx license is not valid"
- alert: QHxManagerReconcileErrors
expr: rate(qhx_manager_reconcile_total{result="error"}[5m]) > 0.1
for: 2m
labels:
severity: warning
annotations:
summary: "Manager reconcile error rate elevated for {{ $labels.controller }}"
- alert: QHxAgentWorkloadAPIDown
expr: qhx_agent_workload_api_connected == 0
for: 1m
labels:
severity: critical
annotations:
summary: "QHx Agent cannot reach the SPIFFE Workload API"
- alert: QHxProxyAuthFailures
expr: rate(qhx_proxy_connections_total{result="auth_failed"}[5m]) > 0.5
for: 2m
labels:
severity: warning
annotations:
summary: "Elevated mTLS authentication failures on {{ $labels.protocol }} listener"

QHx provides reference dashboard templates for:

  • Security operations - Authentication/authorization events, policy violations, threat indicators
  • Component health - Service availability, resource utilization, error rates
  • Performance - Request latency, throughput, cache efficiency
  • Compliance - Audit coverage, policy compliance, vulnerability status

Dashboard templates are available for:

  • Grafana
  • Kibana
  • Splunk

Typical metrics footprint:

  • Metrics per component: ~200-500 time series
  • Scrape interval: 30 seconds (recommended)
  • Retention: 15 days minimum (30 days recommended)

Typical log footprint:

  • Log volume: ~100KB/day per component (INFO level)
  • Security events: ~10KB/day per component
  • Retention: 90 days minimum (1 year recommended)

The Manager metrics endpoint defaults to HTTPS with authentication (--metrics-secure=true). For in-cluster Prometheus scraping without client-cert setup, set --metrics-secure=false in dev clusters.

Restrict Prometheus scrape access with a NetworkPolicy that only allows ingress from the monitoring namespace:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: qhx-metrics-scrape
namespace: qhx-system
spec:
podSelector:
matchLabels:
app.kubernetes.io/part-of: qhx
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: monitoring
ports:
- protocol: TCP
port: 9090
- protocol: TCP
port: 9091
- protocol: TCP
port: 9092

Logs may contain sensitive information:

  • SPIFFE IDs
  • Namespace and service account names
  • User identifiers
  • Failure reasons

Implement RBAC for log access:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: qhx-log-reader
namespace: qhx-system
rules:
- apiGroups: [""]
resources: ["pods/log"]
verbs: ["get", "list"]

Consider filtering sensitive fields before external shipping:

# Fluent Bit filter to remove sensitive fields
[FILTER]
Name modify
Match kube.qhx.*
Remove user_groups
Remove source_ip

QHx telemetry supports compliance requirements:

  • AU-2 (Audit Events) - Security-relevant events logged
  • AU-3 (Content of Audit Records) - Structured event data with context
  • AU-6 (Audit Review) - Integration with SIEM/log analysis tools
  • AU-12 (Audit Generation) - Automated event generation
  • SI-4 (System Monitoring) - Continuous monitoring capabilities

QHx telemetry addresses STIG requirements for:

  • Security event logging (V-230224)
  • Audit log protection (V-230225)
  • Monitoring and alerting (V-230226)
  1. Deploy ServiceMonitor:

    Terminal window
    kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yaml
  2. Verify metrics collection:

    Terminal window
    kubectl port-forward -n monitoring svc/prometheus 9090:9090
    # Open http://localhost:9090/targets and look for qhx-* entries
  3. Scrape a component directly:

    Terminal window
    # Manager (HTTP — set --metrics-secure=false in dev)
    kubectl port-forward -n qhx-system deploy/manager 9090:9090
    curl http://localhost:9090/metrics | grep qhx_manager
    # Agent
    kubectl port-forward -n qhx-system daemonset/qhx-agent 9091:9091
    curl http://localhost:9091/metrics | grep qhx_agent
  1. Deploy Fluent Bit DaemonSet:

    Terminal window
    kubectl apply -f fluent-bit-daemonset.yaml
  2. Configure log parsing:

    Terminal window
    kubectl apply -f fluent-bit-config.yaml
  3. Verify log collection:

    Terminal window
    # Query Elasticsearch
    curl -X GET "localhost:9200/qhx-logs-*/_search?pretty"

For detailed metric specifications, alert rules, and integration examples:

Contact your Messier 42 account team for:

  • Complete metric reference
  • Sample alert rules
  • Dashboard templates
  • Integration guides

Access detailed documentation in the customer portal:

  • Complete telemetry reference
  • CSRMC integration guide
  • Alert rule library
  • Troubleshooting procedures

Verify metrics collection:

Terminal window
# Check Prometheus targets
kubectl exec -n monitoring prometheus-0 -- promtool query instant 'up{job=~"qhx.*"}'

Verify log collection:

Terminal window
# Check recent logs
kubectl logs -n qhx-system -l app.kubernetes.io/name=qhx --tail=100

Metrics not appearing:

  1. Check ServiceMonitor configuration
  2. Verify network policy allows Prometheus scraping
  3. Confirm metrics endpoint is accessible:
    Terminal window
    kubectl port-forward -n qhx-system deploy/manager 9090:9090
    curl http://localhost:9090/metrics

Logs not appearing:

  1. Check Fluent Bit DaemonSet is running
  2. Verify log path matches QHx containers
  3. Check Fluent Bit configuration:
    Terminal window
    kubectl logs -n logging -l app=fluent-bit

Metrics:

  • Short-term (high resolution): 15-30 days
  • Long-term (downsampled): 1 year
  • Use Thanos or Cortex for long-term storage

Logs:

  • Hot storage: 90 days
  • Warm storage: 1 year
  • Cold storage (archive): 7 years (compliance requirements)
  • Start with critical alerts only
  • Tune thresholds based on baseline
  • Implement alert aggregation
  • Use alert routing and silencing

Metrics:

  • Adjust scrape interval based on needs (15s-60s)
  • Use metric relabeling to drop unnecessary labels
  • Implement federation for large deployments

Logs:

  • Use log sampling for high-volume events
  • Implement log level filtering (DEBUG → INFO in production)
  • Use log aggregation before shipping