Skip to content

Capacity Planning and Scaling

QHx capacity planning involves understanding how your deployment topology, workload distribution, and resource allocation affect performance and availability. This guide is based on proven patterns adapted for QHx.

Key planning areas:

  • Deployment topology selection (single, nested, federated)
  • PKI Server sizing based on workload count
  • High availability configuration
  • Datastore performance optimization

Target audience: Platform architects, SREs, capacity planners

Performance factors:

  1. Number of registration entries - More entries = more memory/CPU
  2. SVID TTL - Shorter TTL = more frequent rotation = higher load
  3. Number of agents - Each agent syncs every 5 seconds
  4. Workload distribution - Dense nodes increase agent load
  5. Datastore performance - Often the biggest bottleneck

PKI Server resource consumption:

  • Memory and CPU grow proportionally to registration entries
  • Authorization checks happen on every agent sync (every 5 seconds)
  • Datastore queries are relatively expensive

┌─────────────────────────────────────────────────┐
│ Trust Domain: qhx.dev │
│ │
│ ┌──────────────────────────────────────────┐ │
│ │ PKI Server Cluster (HA) │ │
│ │ ┌─────────┐ ┌─────────┐ ┌─────────┐ │ │
│ │ │ Server1 │ │ Server2 │ │ Server3 │ │ │
│ │ └────┬────┘ └────┬────┘ └────┬────┘ │ │
│ │ └───────────┬──────────────┘ │ │
│ │ │ │ │
│ │ ┌────────▼────────┐ │ │
│ │ │ Shared Datastore│ │ │
│ │ │ (PostgreSQL) │ │ │
│ │ └─────────────────┘ │ │
│ └──────────────────────────────────────────┘ │
└─────────────────────────────────────────────────┘

When to use:

  • <5,000 workloads
  • Single cloud provider, single region
  • Simple administrative domain

Advantages: Simplest to manage, single CA Disadvantages: Datastore bottleneck, limited geographic distribution


Section titled “Topology 2: Nested QHx (Recommended for Multi-Region)”
┌──────────────────────────────────────────────────────────┐
│ Top-Level (Global) PKI Server │
│ Root CA │
└─────────┬─────────────────────────────────────────────────┘
│ Issues intermediate CAs
┌─────┴──────┬──────────────┬─────────────┐
│ │ │ │
┌───▼────┐ ┌───▼────┐ ┌───▼────┐ ┌───▼────┐
│US-EAST │ │US-WEST │ │ EU │ │ AP-SE │
│Regional│ │Regional│ │Regional│ │Regional│
│PKI Srv │ │PKI Srv │ │PKI Srv │ │PKI Srv │
│Own DB │ │Own DB │ │Own DB │ │Own DB │
└────────┘ └────────┘ └────────┘ └────────┘

When to use:

  • 5,000-100,000+ workloads
  • Multi-region or multi-cloud
  • Datastore becoming bottleneck

Advantages:

  • Each region has its own datastore (eliminates cross-region DB)
  • Failure domains isolated
  • Better performance (regional servers closer to workloads)

Configuration:

# Regional server (US-EAST)
spec:
pkiServer:
tier: intermediate
upstreamAuthority:
spire:
socketPath: /run/spire/sockets/agent.sock
datastore:
type: postgres
connectionString: "postgresql://us-east-db..."

When to use:

  • Multiple administrative domains
  • Separate staging/production
  • Regulatory requirements (data sovereignty)

See: Federation Setup Guide


Order-of-magnitude guidelines based on test environments:

Workloads10 Agents100 Agents1,000 Agents5,000 Agents
102 Servers
1 CPU, 1GB
2 Servers
2 CPU, 2GB
2 Servers
4 CPU, 4GB
2 Servers
8 CPU, 8GB
1002 Servers
2 CPU, 2GB
2 Servers
2 CPU, 2GB
2 Servers
8 CPU, 8GB
2 Servers
16 CPU, 16GB
1,0002 Servers
16 CPU, 8GB
2 Servers
16 CPU, 8GB
2 Servers
16 CPU, 8GB
4 Servers
16 CPU, 8GB
10,0004 Servers
16 CPU, 16GB
4 Servers
16 CPU, 16GB
4 Servers
16 CPU, 16GB
8 Servers
16 CPU, 16GB

Notes:

  • “2 Servers” = High availability
  • Does not include datastore sizing
  • Assumes 1-hour SVID TTL

Requirements:

  1. Shared datastore (all servers)
  2. Minimum 2 servers (recommended 3)
  3. LoadBalancer for agent connections
apiVersion: apps/v1
kind: StatefulSet
spec:
replicas: 3 # HA
template:
spec:
containers:
- name: pki-server
resources:
requests:
cpu: 4000m
memory: 4Gi

Critical: Datastore is Often the Bottleneck

Section titled “Critical: Datastore is Often the Bottleneck”

Problem: Authorization checks (every 5 seconds per agent) are expensive queries.

Solutions:

  1. Use PostgreSQL (not SQLite) for production
  2. Tune PostgreSQL settings
  3. Use nested topology for >10,000 workloads
-- /etc/postgresql/postgresql.conf
max_connections = 200
shared_buffers = 4GB
effective_cache_size = 12GB
work_mem = 64MB
random_page_cost = 1.1 # For SSD

Indexes:

CREATE INDEX idx_registered_entries_spiffe_id ON registered_entries(spiffe_id);
CREATE INDEX idx_registered_entries_parent_id ON registered_entries(parent_id);

When:

  • CPU >70% on all servers
  • Agent connection errors

How:

Terminal window
kubectl -n qhx-system scale statefulset pki-server --replicas=5

When:

  • Memory >80%
  • Single server bottleneck

How:

resources:
requests:
cpu: 8000m # Was 4000m
memory: 16Gi # Was 8Gi

When:

  • 10,000 workloads

  • Datastore bottleneck
  • Multi-region

Benefits:

  • Each regional DB handles smaller subset
  • Authorization load distributed
  • Regional failures isolated

  • Estimate workloads (current + 12-month projection)
  • Determine topology (single vs nested vs federated)
  • Select datastore (PostgreSQL for production)
  • Plan for HA (minimum 2 servers)
  • Use sizing table as starting point
  • Account for 2-3x growth
  • Size datastore separately
  • Plan for burst capacity
  • Deploy with monitoring
  • Set up capacity alerts
  • Test failover scenarios
  • Review resource usage monthly
  • Monitor registration entry growth
  • Tune datastore as needed

# PKI Server CPU
rate(container_cpu_usage_seconds_total{pod=~"pki-server.*"}[5m])
# PKI Server memory
container_memory_usage_bytes{pod=~"pki-server.*"}
# Registration entries
spire_server_registration_entries_count
# Datastore query latency
histogram_quantile(0.95, rate(spire_server_datastore_operation_duration_seconds_bucket[5m]))
- alert: PKIServerCPUHigh
expr: rate(container_cpu_usage_seconds_total{pod=~"pki-server.*"}[5m]) > 0.8
annotations:
summary: "Consider scaling PKI Servers"
- alert: DatastoreLatencyHigh
expr: histogram_quantile(0.95, rate(spire_server_datastore_operation_duration_seconds_bucket[5m])) > 1.0
annotations:
summary: "Datastore slow - tune or scale"

Key decisions:

  • <5,000 workloads: Single trust domain
  • 5,000-100,000: Nested topology
  • Multiple admin domains: Federated

Critical: Datastore performance is typically the bottleneck—plan carefully!