High availability and failure domains
Replicate only supported components, remove avoidable single points of failure, and test disruption against explicit objectives.
Edge routing, database, storage, scheduling capacity, and leader behavior must survive the same failure. Merely running two copies without testing the loss of one domain leaves you with two identical failures, not redundancy.
Define the supported contract
Publish replica safety, leader election, migration ownership, session behavior, and remaining single points of failure.
Before deploying with redundancy, decide which components will replicate and at what minimum count. The operator replicates safely at 2+ replicas, but the API must stay at 1 replica; game servers themselves are stateless replicas (one per instance) and scale independently.
- Operator component: Leader-elected across replicas. Safe at 2+ replicas with
operator.replicasin Helm values. Leader election uses Kubernetes leases; the operator exits on missed lease renewal, triggering a pod restart. A single-node cluster with one operator replica is fault-tolerant only within that node. - API component: Keep
api.replicas: 1with either database driver: SQLite (default, production-tested) is single-writer on a RWO volume, and with PostgreSQL (experimental) the user-management lock and audit hash chain are still per-process, so extra replicas are not safe. Game agents talk to the API via Kubernetes Service discovery and automatically reconnect on pod restarts. The API stores game state, audit logs, and administrative data; loss of the API pod does not interrupt active game servers (they remain running and playable). - Game server pods: Each GameServer pod is stateless application logic (plus ephemeral logs). World data persists on a PVC backed by your StorageClass. Multiple GameServers are independent replicas; loss of one pod does not affect others. Game servers are typically single-replica per instance; replication happens at the instance level, not the pod level.
- Ingress controller, load balancer, and DNS: Provided by your Kubernetes distribution and network provider. Gameplane does not manage these; your infrastructure must provide redundancy (e.g., multi-zone load balancer, multiple ingress replicas, DNS failover).
State database requirements and constraints:
- SQLite (default): Single writer on a RWO PVC. Database file is mutable in place; concurrent writers corrupt it. Backups via restic snapshots. Simpler operational model; suitable for single-operator setups and development.
- PostgreSQL (experimental): An external, separately administered database. Replication, point-in-time recovery, and connection pooling are PostgreSQL-native concerns. Not yet recommended for production: the published api image is built without it, it is not yet covered by e2e or upgrade tests, and it does not make multiple API replicas safe.
Separate cluster-scoped dependencies from application replicas:
- The operator, API, and sentinel (wake-on-connect) are control-plane components that manage the cluster. They have their own replicas and heartbeat logic.
- Game server pods are workload instances. A crashed game pod is restarted by the operator; a crashed control-plane pod is restarted by Kubernetes. They have different failure modes and recovery times.
Spread dependencies
Place safe replicas and durable dependencies across nodes, zones, and independent failure domains.
Running 2+ operator replicas on the same node does not prevent node loss. Spread replicas across failure domains using pod anti-affinity, topology spread constraints, and disruption budgets. Test that loss of any single domain (node, availability zone, data center) leaves you with:
- Operator can still elect a leader and reconcile changes
- API database and data remain available
- Ingress routing and load balancer continue serving game traffic
- Backup and storage systems are reachable
Use anti-affinity, topology spread, and appropriate disruption budgets:
- Anti-affinity: Prefer (soft) or require (hard) that replicas land on different nodes. Hard anti-affinity can cause pod pending if insufficient nodes exist. Example:
podAntiAffinity.preferredDuringSchedulingIgnoredDuringExecutionspreads operator replicas. - Topology spread: Distribute pods across topology keys such as
kubernetes.io/hostname(nodes),topology.kubernetes.io/zone(availability zones), ortopology.kubernetes.io/region(regions). Helm values or overlay patches apply these to system components. - Pod Disruption Budgets (PDB): Prevent simultaneous eviction of too many replicas during cluster maintenance (node drain, upgrades). Set a minimum of 1 available for control-plane components (operator, API) so the cluster always has capacity to reschedule.
Provide redundant edge, database, storage, and backup paths:
- Ingress and load balancer: Multiple ingress replicas across zones; load balancer with health checks and failover. Your cloud provider or on-premises LB provides this.
- Database (API SQLite or PostgreSQL): Backup to a geographically separate location (S3 bucket in a different region, off-site restic repository). Test restore time.
- Storage (PVCs for game data): Use a StorageClass backed by replicated storage (CEPH, cloud provider block storage with snapshots, etc.). Single-disk NAS is not HA. If your storage is single-node, loss of that node loses all game worlds until backup restoration.
- DNS, certificates, and image registry: Multi-zone DNS (managed by your cloud provider or external service). Certificates from a provider with redundancy (Let’s Encrypt, your corporate CA with replication). Image registry with geo-replication or a local registry mirror.
Reserve failover capacity instead of saturating every node:
- Do not schedule game pods with resource requests that consume 100% of node capacity. Reserve 10–20% headroom for operator, API, and agent sidecars, plus 10–20% for Kubernetes system daemons. On a node with 8 CPUs, schedule no more than 6 CPUs of game workload.
- When a node is drained (e.g., during maintenance), game pods evict and reschedule to other nodes. If no capacity exists, pods remain pending. Oversub provisioning capacity beforehand.
- Pod Disruption Budgets (PDB) prevent simultaneous disruption of control-plane components, but also cause scheduling to fail if eviction would breach the budget. For example, if
minAvailable: 1on a 2-replica operator and you drain a node with one operator pod, the eviction is allowed (the other replica survives). But if you drain a node with both replicas, eviction fails and the node drain blocks until pods are manually deleted or the budget is relaxed.
Test disruption
Drain or terminate one dependency at a time and measure availability, reconciliation continuity, handoff, session behavior, and audit trail.
A system that is never tested for failure will fail in unexpected ways. Plan a maintenance window and execute controlled disruption tests. Record the results (RTO, RPO, error rates, session drops) so you have a baseline and confidence in your setup.
Drain or terminate one dependency at a time and measure:
- Availability: Is the dashboard or game API reachable during the outage? Do requests queue, fail, or time out?
- Reconciliation continuity: Did the operator fully process pending GameServer changes, or are any stuck in a transitional state?
- Handoff and session behavior: If an API pod crashed, did agents re-authenticate? Did WebSocket clients reconnect?
- Jobs and backups: In-flight Backup or Restore jobs that lose their pod must retry from the API. Do they recover?
- Audit trail: Are the events recorded in the audit log, or did events drop during the outage?
HA TEST
Common setup patterns
Single-node development cluster:
- 1 operator replica, 1 API replica (SQLite).
- No anti-affinity, no multi-zone.
- Sufficient for testing; not HA.
Multi-node HA cluster:
- 2–3 operator replicas spread across nodes (anti-affinity).
- A single API replica; PostgreSQL (experimental) is optional and does not make extra API replicas safe.
- StorageClass backed by replicated storage (CEPH, cloud provider block, etc.).
- 2+ ingress replicas, multi-zone load balancer.
- Off-site backup destination.
Geographic separation (multi-cluster, stretch cluster):
- Either run separate Gameplane control planes in each region or data center, or run one central API and dashboard and register each region’s cluster, with its own operator and an optional agent gateway (v0.3.0). A gateway outage does not stop games or their local operator.
- Game servers stay local to their data center; they do not migrate.
- Backup and restore allow recovery to a different cluster if a whole region fails.
See Multi-cluster topology for cross-region failover patterns, and Platform disaster recovery for restoring after catastrophic loss.