High Availability with PostgreSQL
QHx supports running SPIRE Server in high-availability (HA) mode with multiple replicas sharing an external PostgreSQL datastore. This guide covers both in-cluster PostgreSQL operators and managed services such as AWS RDS.
Prerequisites
Section titled “Prerequisites”- QHx installed in your cluster via Helm (see Helm-Based Installation)
- A
StorageClassthat supportsReadWriteOncefor SPIRE Server PVCs kubectlconfigured for your cluster- Customer credentials for
oci.messier42.com
Step 1 — Deploy a PostgreSQL Operator
Section titled “Step 1 — Deploy a PostgreSQL Operator”Amazon RDS for PostgreSQL is a managed option that removes the operational burden of running PostgreSQL inside the cluster.
Create the RDS instance (or use an existing one) and note the endpoint,
e.g. spire.cxxx.us-east-1.rds.amazonaws.com.
Create the spire database and user:
CREATE DATABASE spire;CREATE USER spire WITH PASSWORD '<strong-password>';GRANT ALL PRIVILEGES ON DATABASE spire TO spire;Ensure the RDS security group allows inbound TCP on port 5432 from the CIDR or security group of your cluster nodes.
The DSN to use in Step 2:
postgresql://spire:<password>@spire.cxxx.us-east-1.rds.amazonaws.com:5432/spire?sslmode=requireRDS uses a certificate from Amazon’s CA. If your cluster doesn’t have that CA
in its trust store, add sslrootcert=/path/to/aws-rds-ca.pem to the DSN, or
use sslmode=require without certificate verification (acceptable for most
in-cluster use cases):
postgresql://spire:<password>@spire.cxxx.us-east-1.rds.amazonaws.com:5432/spire?sslmode=requireYou do not need to deploy a PostgreSQL operator — skip straight to Step 2 once the RDS instance is ready.
CloudNativePG is a CNCF project and the recommended option for production clusters.
kubectl apply --server-side -f \ https://raw.githubusercontent.com/cloudnative-pg/cloudnative-pg/release-1.24/releases/cnpg-1.24.0.yamlWait for the operator to be ready:
kubectl rollout status deployment/cnpg-controller-manager -n cnpg-systemCreate a PostgreSQL cluster with a dedicated spire database:
apiVersion: postgresql.cnpg.io/v1kind: Clustermetadata: name: spire-pg namespace: qhx-systemspec: instances: 3 storage: size: 5Gi bootstrap: initdb: database: spire owner: spire secret: name: spire-pg-credentialsCreate the password secret first:
kubectl create secret generic spire-pg-credentials \ --namespace qhx-system \ --from-literal=username=spire \ --from-literal=password='<strong-password>'Then apply the cluster manifest:
kubectl apply -f spire-pg-cluster.yamlWait for all instances to be ready:
kubectl wait cluster/spire-pg -n qhx-system \ --for=condition=Ready --timeout=5mThe CloudNativePG operator creates a read-write service named
spire-pg-rw.<namespace>.svc.cluster.local that SPIRE Server connects to.
The Bitnami PostgreSQL Helm chart provides a simpler single-primary setup suitable for smaller clusters.
helm repo add bitnami https://charts.bitnami.com/bitnamihelm repo update
helm install spire-pg bitnami/postgresql \ --namespace qhx-system \ --set auth.database=spire \ --set auth.username=spire \ --set auth.password='<strong-password>' \ --set primary.persistence.size=5GiThe service is available at spire-pg-postgresql.qhx-system.svc.cluster.local.
Step 2 — Store the DSN in a Secret
Section titled “Step 2 — Store the DSN in a Secret”SPIRE Server connects to PostgreSQL using the pkiDatastoreConnectionString
field. Because this value contains credentials it must be stored in the
manager Secret, not the ConfigMap.
The Secret must use the key qhx-manager-secret.yaml and contain valid YAML:
pkiDatastoreConnectionString: "postgresql://spire:<password>@spire-pg-rw.qhx-system.svc.cluster.local:5432/spire?sslmode=require"Create or update the Secret in the cluster:
kubectl create secret generic qhx-manager-secret \ --namespace qhx-system \ --from-file=qhx-manager-secret.yaml=./qhx-manager-secret.yaml \ --dry-run=client -o yaml | kubectl apply -f -Step 3 — Configure the Manager Secret Reference
Section titled “Step 3 — Configure the Manager Secret Reference”If you did not set managerSecretName during the initial install, upgrade the
release to add it now. Use the same version and credentials you used when
installing QHx:
$ helm upgrade qhx oci://oci.messier42.com/qhx/charts/qhx-core \ --version <INSTALLED-VERSION> \ --namespace qhx-system \ --reuse-values \ --set managerSecretName=qhx-manager-secret \ --set managerSecretNamespace=qhx-system \ --waitStep 4 — Enable HA in the ConfigMap
Section titled “Step 4 — Enable HA in the ConfigMap”Edit the qhx-manager ConfigMap to enable multiple SPIRE Server replicas and
configure storage for the per-replica PVCs:
kubectl edit configmap qhx-manager -n qhx-systemAdd or update these fields:
pkiServerReplicas: "2" # number of SPIRE Server replicas (≥ 2 for HA)pkiStorageClassName: "standard" # StorageClass that supports ReadWriteOncepkiStorageSize: "1Gi" # size of each replica's PVCThe manager watches the ConfigMap and picks up changes without a restart. It will trigger a rolling update of the SPIRE Server StatefulSet automatically.
Step 5 — Verify the HA Deployment
Section titled “Step 5 — Verify the HA Deployment”Check that all SPIRE Server pods are running:
kubectl get pods -n qhx-system -l app.kubernetes.io/component=pki-serverExpected output (2-replica example):
NAME READY STATUS RESTARTS AGEqhx-spire-server-0 1/1 Running 0 2mqhx-spire-server-1 1/1 Running 0 90sVerify that one PVC was created per replica:
kubectl get pvc -n qhx-system -l app.kubernetes.io/component=pki-serverConfirm that SPIRE Server logs show the PostgreSQL datastore plugin:
kubectl logs -n qhx-system qhx-spire-server-0 | grep -i datastoreYou should see lines similar to:
INFO Datastore connected subsystem_name=catalog type=sql driver=postgresBehaviour Notes
Section titled “Behaviour Notes”| Aspect | Detail |
|---|---|
| Pod placement | The manager configures a soft (PreferredDuringScheduling) anti-affinity rule so replicas prefer different nodes. |
| Startup order | ParallelPodManagement is enabled — all pods start simultaneously rather than one-by-one. |
| Leader election | SPIRE Server leader election is automatically enabled when pkiServerReplicas > 1. |
| PVC lifecycle | Each replica gets its own PVC via VolumeClaimTemplates. Scaling down does not delete PVCs automatically. |
| ConfigMap / Secret reload | Both are watched live; no manager restart is required when you change the DSN or replica count. |
| StatefulSet VolumeClaimTemplates | These are immutable in Kubernetes. If you need to change pkiStorageClassName or pkiStorageSize after the StatefulSet exists, the manager will delete and recreate it automatically. |
Rollback to Single-Replica Mode
Section titled “Rollback to Single-Replica Mode”To disable HA and revert to the embedded SQLite datastore:
- Set
pkiServerReplicas: "1"in the ConfigMap. - Remove or clear
pkiDatastoreConnectionStringfrom the Secret. - The manager will update the StatefulSet; the unused PVCs can be deleted manually.