Monitoring and Observability
Overview
Section titled “Overview”QHx provides comprehensive telemetry for security operations centers, enabling integration with enterprise monitoring systems and DoD cybersecurity operations infrastructure. All QHx components expose structured metrics and logs that integrate seamlessly with standard observability tools.
Target audience: Platform operators, security operations teams, system administrators.
Telemetry Outputs
Section titled “Telemetry Outputs”Metrics
Section titled “Metrics”QHx components expose Prometheus-compatible metrics for:
- Authentication events - Node and workload attestation success/failure rates
- Authorization events - Policy enforcement decisions and denials
- Cryptographic operations - Certificate issuance, rotation, and lifecycle events
- Component health - Service availability, resource utilization, and performance
- Operational metrics - Request rates, latency, and throughput
Supported metric collectors:
- Prometheus (recommended)
- StatsD / DogStatsD
- M3
- OpenTelemetry Collector
Structured Logging
Section titled “Structured Logging”QHx components emit structured JSON logs to stdout/stderr, enabling integration with standard Kubernetes logging stacks:
- Security events - Authentication failures, authorization denials, policy violations
- Operational events - Component lifecycle, configuration changes, errors
- Audit events - Administrative actions, policy modifications, access attempts
Log fields include:
- Timestamp (RFC3339)
- Severity level
- Component identifier
- Event type and category
- Contextual metadata (namespace, user, SPIFFE ID)
Supported log aggregators:
- Fluent Bit / Fluentd (recommended)
- Logstash
- OpenTelemetry Collector
- Splunk
- Datadog
DoD Integration
Section titled “DoD Integration”QHx supports integration with DoD cybersecurity operations infrastructure:
CSRMC (Cybersecurity Service Resource Management Center)
Section titled “CSRMC (Cybersecurity Service Resource Management Center)”QHx telemetry is designed for CSRMC integration, providing:
- Security event correlation
- Incident detection and alerting
- Compliance monitoring
- Situational awareness
ACAS (Assured Compliance Assessment Solution)
Section titled “ACAS (Assured Compliance Assessment Solution)”QHx supports ACAS vulnerability scanning:
- Container image scanning
- Configuration assessment
- TLS/SSL certificate validation
HBSS (Host-Based Security System)
Section titled “HBSS (Host-Based Security System)”QHx telemetry can feed HBSS for:
- File integrity monitoring
- Process monitoring
- Network connection monitoring
Metrics Collection
Section titled “Metrics Collection”Every QHx control-plane component exposes a /metrics endpoint on a dedicated
port. All endpoints use the prometheus named port and carry the label
qhx.dev/prometheus-scrape: "true" on their Kubernetes Services.
| Component | Deployed Port | Flag / Config |
|---|---|---|
| SPIRE server | 9998 | SPIRE telemetry block |
| SPIRE agent | configurable | SPIRE telemetry block |
| Manager | 9090 | --metrics-bind-address |
| Agent | 9091 | --metrics-addr |
| Proxy | 9092 | metricsAddr (config) |
The Manager endpoint is plain HTTP (--metrics-secure=false in the cluster
deployment). Set any address flag to 0 or "" to disable that component’s
endpoint.
Prometheus Integration
Section titled “Prometheus Integration”Apply the existing ServiceMonitor — it selects on qhx.dev/prometheus-scrape: "true"
with a 30 s scrape interval and covers SPIRE instances, Manager, and Agent:
kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yamlThe Proxy runs as an injected sidecar, so its /metrics endpoint appears on
user pods rather than a dedicated Service. Use a PodMonitor to scrape it:
apiVersion: monitoring.coreos.com/v1kind: PodMonitormetadata: name: qhx-proxy namespace: monitoringspec: namespaceSelector: any: true selector: matchLabels: qhx.dev/proxy-injected: "true" podMetricsEndpoints: - port: qhx-proxy-metrics path: /metrics interval: 30sHelm values:
manager: metricsAddr: ":9090" # --metrics-bind-address metricsSecure: false # --metrics-secure (set false for plain HTTP)
agent: metricsAddr: ":9091"
proxy: metricsAddr: ":9092"Alternative Collectors
Section titled “Alternative Collectors”StatsD configuration:
# PKI Server configurationtelemetry { Prometheus { enabled = false } Statsd { enabled = true address = "statsd-agent.monitoring:8125" }}Metrics Reference
Section titled “Metrics Reference”Manager Metrics
Section titled “Manager Metrics”Reconciliation
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_reconcile_total | Counter | controller, result | Reconcile attempts. controller: spire-instance | qhx-cluster-policy | bundle-exchange | license-status. result: success | error | requeue |
qhx_manager_reconcile_duration_seconds | Histogram | controller | Wall time of each reconcile call |
qhx_manager_reconcile_requeue_after_seconds | Histogram | controller | Requested requeue delay when result is requeue |
SPIRE Instance State
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_spire_instances_total | Gauge | — | Total desired SPIRE instances after policy evaluation |
qhx_manager_spire_instances_ready | Gauge | — | SPIRE instances with at least one ready agent pod on the manager node |
qhx_manager_spire_instances_deleted_total | Counter | — | Instances deleted after cooldown expiry or explicit removal |
qhx_manager_spire_instances_pvc_rebuild_total | Counter | — | StatefulSet delete+recreate cycles triggered by immutable VolumeClaimTemplate changes |
qhx_manager_spire_instance_hash_collisions_total | Counter | — | Deduplication events: two policies resolving to the same descriptor hash |
qhx_manager_agent_sources_sync_errors_total | Counter | — | Failures syncing agent sources from the SPIFFE Workload API socket (non-fatal — reconcile proceeds) |
Policy, Webhook, and Bundle Exchange
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_policies_total | Gauge | kind | Policy object count. kind: cluster | namespace |
qhx_manager_webhook_requests_total | Counter | operation, resource, result | Admission webhook calls. result: allowed | denied | error |
qhx_manager_webhook_duration_seconds | Histogram | operation, resource | Admission webhook latency |
qhx_manager_bundle_exchange_total | Counter | direction, result | Bundle exchange operations. direction: push | pull. result: success | error |
License and Config
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_manager_license_valid | Gauge | — | 1 if at least one valid license is present, 0 otherwise |
qhx_manager_config_reloads_total | Counter | result, source | ConfigMap/Secret reload events. result: success | error. source: configmap | secret (identifies what failed) | all (on success) |
qhx_manager_config_watch_errors_total | Counter | source | Watch loop failures (auto-restart after 5 s). source: configmap | secret |
Agent Metrics
Section titled “Agent Metrics”| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_agent_workload_api_connected | Gauge | — | 1 when the SPIFFE Workload API socket is reachable, 0 after shutdown |
qhx_agent_svid_renewals_total | Counter | result | SVID renewal attempts. result: success | error |
qhx_agent_svid_expiry_seconds | Gauge | spiffe_id | Seconds until the current SVID for the given SPIFFE ID expires |
qhx_agent_workload_api_watch_errors_total | Counter | — | Errors from the Workload API watch stream |
Proxy Metrics
Section titled “Proxy Metrics”Connections and Traffic
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_connections_total | Counter | protocol, direction, result | Connection attempts. protocol: http | tcp | mqtt. direction: inbound | outbound. result: success | auth_failed | error |
qhx_proxy_active_connections | Gauge | protocol, direction | Currently open connections |
qhx_proxy_bytes_received_total | Counter | protocol | Bytes received from clients |
qhx_proxy_bytes_sent_total | Counter | protocol | Bytes forwarded to backends |
Latency
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_tls_handshake_duration_seconds | Histogram | protocol, role | mTLS handshake time including SPIFFE SVID validation. role: server | client. Currently only emitted for mqtt; HTTP and TCP delegate TLS to the standard library |
qhx_proxy_request_duration_seconds | Histogram | protocol, http_method | End-to-end request latency (HTTP only) |
Authentication and Notary
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_auth_decisions_total | Counter | result, reason, protocol, role | SPIFFE identity auth outcomes. result: allowed | denied. reason: svid_invalid | policy_deny | ok. protocol: http | tcp | mqtt. role: server | client |
qhx_proxy_notary_records_total | Counter | result, protocol | Audit records written by the notary middleware. result: success | error. protocol: http | mqtt |
qhx_proxy_notary_level_total | Counter | level | Notarization level distribution. level: workload | logRequest | signRequest |
qhx_proxy_notary_db_errors_total | Counter | op, entity | bbolt database errors. op: put | get | init | close. entity: certificate_blob | workload_statement | receipt | db |
MQTT Buffer
These metrics are only populated when the MQTT protocol is in use.
| Metric | Type | Labels | Description |
|---|---|---|---|
qhx_proxy_mqtt_buffer_messages_queued | Gauge | — | Messages currently held in the store-and-forward buffer (not yet populated — requires middleware-layer instrumentation) |
qhx_proxy_mqtt_buffer_messages_stored_total | Counter | — | Messages successfully forwarded to the upstream publisher |
qhx_proxy_mqtt_buffer_messages_acked_total | Counter | source | Messages for which a PUBACK was sent to the client. source: middleware (buffer layer handled the publish) | upstream (forwarded to broker and confirmed) |
qhx_proxy_mqtt_buffer_replay_failures_total | Counter | — | Times the replay loop exhausted its backoff budget (not yet populated — requires middleware-layer instrumentation) |
qhx_proxy_mqtt_upstream_connack_rejected_total | Counter | reason | Upstream CONNACK rejections by reason code |
qhx_proxy_mqtt_upstream_reconnections_total | Counter | result | Upstream reconnection attempts. result: success | error |
Log Collection
Section titled “Log Collection”Fluent Bit Integration
Section titled “Fluent Bit Integration”ConfigMap example:
apiVersion: v1kind: ConfigMapmetadata: name: fluent-bit-config namespace: loggingdata: fluent-bit.conf: | [INPUT] Name tail Path /var/log/containers/qhx-*_qhx-system_*.log Parser docker Tag kube.qhx.* Refresh_Interval 5 Mem_Buf_Limit 5MB Skip_Long_Lines On
[FILTER] Name parser Match kube.qhx.* Key_Name log Parser json Reserve_Data On
[FILTER] Name modify Match kube.qhx.* Add cluster_name ${CLUSTER_NAME} Add environment ${ENVIRONMENT}
[OUTPUT] Name es Match kube.qhx.* Host elasticsearch.monitoring Port 9200 Index qhx-logs Logstash_Format On Logstash_Prefix qhxAlerting
Section titled “Alerting”QHx telemetry supports integration with alerting systems:
- Prometheus AlertManager
- Grafana
- PagerDuty
- Splunk
- ServiceNow
Sample alert rules:
groups:- name: qhx-critical rules: - alert: QHxLicenseInvalid expr: qhx_manager_license_valid == 0 for: 5m labels: severity: critical annotations: summary: "QHx license is not valid"
- alert: QHxManagerReconcileErrors expr: rate(qhx_manager_reconcile_total{result="error"}[5m]) > 0.1 for: 2m labels: severity: warning annotations: summary: "Manager reconcile error rate elevated for {{ $labels.controller }}"
- alert: QHxAgentWorkloadAPIDown expr: qhx_agent_workload_api_connected == 0 for: 1m labels: severity: critical annotations: summary: "QHx Agent cannot reach the SPIFFE Workload API"
- alert: QHxProxyAuthFailures expr: rate(qhx_proxy_connections_total{result="auth_failed"}[5m]) > 0.5 for: 2m labels: severity: warning annotations: summary: "Elevated mTLS authentication failures on {{ $labels.protocol }} listener"Dashboards
Section titled “Dashboards”QHx provides reference dashboard templates for:
- Security operations - Authentication/authorization events, policy violations, threat indicators
- Component health - Service availability, resource utilization, error rates
- Performance - Request latency, throughput, cache efficiency
- Compliance - Audit coverage, policy compliance, vulnerability status
Dashboard templates are available for:
- Grafana
- Kibana
- Splunk
Resource Requirements
Section titled “Resource Requirements”Typical metrics footprint:
- Metrics per component: ~200-500 time series
- Scrape interval: 30 seconds (recommended)
- Retention: 15 days minimum (30 days recommended)
Typical log footprint:
- Log volume: ~100KB/day per component (INFO level)
- Security events: ~10KB/day per component
- Retention: 90 days minimum (1 year recommended)
Security Considerations
Section titled “Security Considerations”Metrics Endpoint Security
Section titled “Metrics Endpoint Security”The Manager metrics endpoint defaults to HTTPS with authentication
(--metrics-secure=true). For in-cluster Prometheus scraping without
client-cert setup, set --metrics-secure=false in dev clusters.
Restrict Prometheus scrape access with a NetworkPolicy that only allows
ingress from the monitoring namespace:
apiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: qhx-metrics-scrape namespace: qhx-systemspec: podSelector: matchLabels: app.kubernetes.io/part-of: qhx ingress: - from: - namespaceSelector: matchLabels: kubernetes.io/metadata.name: monitoring ports: - protocol: TCP port: 9090 - protocol: TCP port: 9091 - protocol: TCP port: 9092Log Access Control
Section titled “Log Access Control”Logs may contain sensitive information:
- SPIFFE IDs
- Namespace and service account names
- User identifiers
- Failure reasons
Implement RBAC for log access:
apiVersion: rbac.authorization.k8s.io/v1kind: Rolemetadata: name: qhx-log-reader namespace: qhx-systemrules:- apiGroups: [""] resources: ["pods/log"] verbs: ["get", "list"]Sensitive Data Filtering
Section titled “Sensitive Data Filtering”Consider filtering sensitive fields before external shipping:
# Fluent Bit filter to remove sensitive fields[FILTER] Name modify Match kube.qhx.* Remove user_groups Remove source_ipCompliance and Audit
Section titled “Compliance and Audit”QHx telemetry supports compliance requirements:
NIST 800-53 Controls
Section titled “NIST 800-53 Controls”- AU-2 (Audit Events) - Security-relevant events logged
- AU-3 (Content of Audit Records) - Structured event data with context
- AU-6 (Audit Review) - Integration with SIEM/log analysis tools
- AU-12 (Audit Generation) - Automated event generation
- SI-4 (System Monitoring) - Continuous monitoring capabilities
STIG Requirements
Section titled “STIG Requirements”QHx telemetry addresses STIG requirements for:
- Security event logging (V-230224)
- Audit log protection (V-230225)
- Monitoring and alerting (V-230226)
Getting Started
Section titled “Getting Started”Basic Prometheus Setup
Section titled “Basic Prometheus Setup”-
Deploy ServiceMonitor:
Terminal window kubectl apply -f doc/examples/m3-demo/spire-servicemonitor.yaml -
Verify metrics collection:
Terminal window kubectl port-forward -n monitoring svc/prometheus 9090:9090# Open http://localhost:9090/targets and look for qhx-* entries -
Scrape a component directly:
Terminal window # Manager (HTTP — set --metrics-secure=false in dev)kubectl port-forward -n qhx-system deploy/manager 9090:9090curl http://localhost:9090/metrics | grep qhx_manager# Agentkubectl port-forward -n qhx-system daemonset/qhx-agent 9091:9091curl http://localhost:9091/metrics | grep qhx_agent
Basic Fluent Bit Setup
Section titled “Basic Fluent Bit Setup”-
Deploy Fluent Bit DaemonSet:
Terminal window kubectl apply -f fluent-bit-daemonset.yaml -
Configure log parsing:
Terminal window kubectl apply -f fluent-bit-config.yaml -
Verify log collection:
Terminal window # Query Elasticsearchcurl -X GET "localhost:9200/qhx-logs-*/_search?pretty"
Detailed Documentation
Section titled “Detailed Documentation”For detailed metric specifications, alert rules, and integration examples:
Evaluation Customers
Section titled “Evaluation Customers”Contact your Messier 42 account team for:
- Complete metric reference
- Sample alert rules
- Dashboard templates
- Integration guides
Existing Customers
Section titled “Existing Customers”Access detailed documentation in the customer portal:
- Complete telemetry reference
- CSRMC integration guide
- Alert rule library
- Troubleshooting procedures
Operational Procedures
Section titled “Operational Procedures”Health Checks
Section titled “Health Checks”Verify metrics collection:
# Check Prometheus targetskubectl exec -n monitoring prometheus-0 -- promtool query instant 'up{job=~"qhx.*"}'Verify log collection:
# Check recent logskubectl logs -n qhx-system -l app.kubernetes.io/name=qhx --tail=100Troubleshooting
Section titled “Troubleshooting”Metrics not appearing:
- Check ServiceMonitor configuration
- Verify network policy allows Prometheus scraping
- Confirm metrics endpoint is accessible:
Terminal window kubectl port-forward -n qhx-system deploy/manager 9090:9090curl http://localhost:9090/metrics
Logs not appearing:
- Check Fluent Bit DaemonSet is running
- Verify log path matches QHx containers
- Check Fluent Bit configuration:
Terminal window kubectl logs -n logging -l app=fluent-bit
Best Practices
Section titled “Best Practices”Retention Policies
Section titled “Retention Policies”Metrics:
- Short-term (high resolution): 15-30 days
- Long-term (downsampled): 1 year
- Use Thanos or Cortex for long-term storage
Logs:
- Hot storage: 90 days
- Warm storage: 1 year
- Cold storage (archive): 7 years (compliance requirements)
Alert Fatigue Reduction
Section titled “Alert Fatigue Reduction”- Start with critical alerts only
- Tune thresholds based on baseline
- Implement alert aggregation
- Use alert routing and silencing
Performance Optimization
Section titled “Performance Optimization”Metrics:
- Adjust scrape interval based on needs (15s-60s)
- Use metric relabeling to drop unnecessary labels
- Implement federation for large deployments
Logs:
- Use log sampling for high-volume events
- Implement log level filtering (DEBUG → INFO in production)
- Use log aggregation before shipping
Related Documentation
Section titled “Related Documentation”- Architecture Overview - System components and design
- Security Hardening - Security configuration
- Deployment Guide - Installation procedures