Platform Disaster Recovery
Rebuild Gameplane on a clean cluster from a tested recovery set, then restore platform identity and worlds in a controlled sequence.
A world backup only captures game data. To recover from a platform-wide outage (cluster loss, database corruption, credential compromise), you must also capture and restore Gameplane’s own state: release versions, configuration, CRDs, database, encryption keys, identity providers, repository credentials, DNS, TLS certificates, and ownership records.
Releases up to and including v0.3.0 are published under ghcr.io/valgulnecron, so this page uses that registry; releases after v0.3.0 move to ghcr.io/gameplanepanel.
Define the recovery set
Capture every versioned artifact, platform state source, credential, trust root, and ownership record required to rebuild.
A disaster recovery manifest should document and preserve:
- Release and configuration: Gameplane version (e.g.,
v0.2.0-beta.8), Helm values (includingapi.db.driver, ingress hostname, OIDC settings, storage class defaults), cluster Kubernetes version, and all base-image versions (restic, busybox, kubectl). - CRD and operator state: All GameServer, GameTemplate, Module, ModuleSource, and Cluster CRD definitions; the operator’s RBAC role and ClusterRole; and the operator’s audit/observability configuration.
- Platform database: SQLite file or PostgreSQL backup dump, including all tables for users, roles, permissions, audit logs, RBAC overrides, and platform settings.
- Secrets and certificates: All platform Secrets (backup repository credentials, OIDC client secrets, API signing keys, TLS certificates, image registry credentials, database credentials, webhook auth tokens), stored securely off-cluster with access controls.
- Module sources and registry credentials: Module source configurations (Git URLs, OCI registries, local paths), registry pull secrets, and Cosign public keys for verification.
- Cluster trust roots: DNS provider credentials (if dynamic), TLS certificate private keys, OIDC IdP certificates, and any custom CA certificates.
- Ownership records: A mapping of GameServer/GameTemplate names to their owners, stored in the database; also any external ownership tracking (e.g., team/permission matrix) kept outside Gameplane.
Before a crisis, rehearse the recovery plan with synthetic or replica data on a staging cluster. A manifest that hasn’t been validated is just a hope.
Rebuild in dependency order
Prepare infrastructure first, install the matching release, then restore platform state before worlds.
A platform-wide rebuild must follow this sequence:
1. Prepare the target Kubernetes cluster
- Ensure the target cluster meets the same specifications as the original: Kubernetes 1.28+, nodes with sufficient capacity for the operator, API, and game workloads, storage backend provisioned, and network connectivity to backup destinations.
- Provision storage: Persistent volumes for API state (if using SQLite), and test access to backup repositories (S3, SFTP, local restic backend).
- Set up edge infrastructure: DNS, TLS certificates, load balancer pools, and any relay services (if using frp/Tailscale for NAT traversal).
- Configure monitoring and audit sinks: Provision the destination for syslog, webhook, or S3 audit events and Prometheus scrape targets.
2. Install the matching Gameplane release
-
Install the exact version that was running when the backup was taken. Use the backup manifest’s recorded Helm values:
helm upgrade --install gameplane oci://ghcr.io/valgulnecron/charts/gameplane \ --version 0.2.0-beta.8 \ --namespace gameplane-system \ --create-namespace \ --values /path/to/recorded-values.yaml -
Verify the operator and API are running:
kubectl -n gameplane-system get deploy -l app.kubernetes.io/name=gameplane kubectl -n gameplane-system logs -f deploy/gameplane-operator -
Do not create GameServers or restore worlds yet. Wait until platform state is restored.
3. Restore the platform database and Secrets
-
Restore the database using the backup dump captured from the original cluster:
- For SQLite: copy the backed-up
gameplane.dbinto the API’s PVC, then restart the API pod. - For PostgreSQL: restore the dump into the database server (ensure the version matches), then verify connectivity from the API pod.
# Example: restore SQLite via kubectl exec kubectl -n gameplane-system exec -i deploy/gameplane-api -- \ cp /dev/stdin /data/gameplane.db < /path/to/gameplane.db.backup kubectl -n gameplane-system rollout restart deploy/gameplane-api - For SQLite: copy the backed-up
-
Restore Secrets that the API and operator reference (OIDC client secrets, backup repo credentials, webhook auth tokens, TLS certs):
kubectl apply -f /path/to/secrets-backup.yamlEnsure Secrets are created in the correct namespaces (
gameplane-systemfor platform Secrets,gameplanefor game-related Secrets). -
Verify the API is healthy and can authenticate users:
kubectl -n gameplane-system logs deploy/gameplane-api | grep "started\|listening" curl -s http://localhost:8080/health # or port-forward if needed
4. Restore CRD definitions and resource ownership
-
Apply the backed-up CRD manifest to register the schema for GameServer, GameTemplate, Module, ModuleSource, Cluster, etc.:
kubectl apply -f /path/to/crds-backup.yaml -
Import backed-up GameTemplates and Clusters (do not yet create or restore GameServers):
kubectl apply -f /path/to/gametemplates-backup.yaml kubectl apply -f /path/to/clusters-backup.yamlThese are lightweight; they define the blueprint and multi-cluster topology without triggering any game workloads.
-
Verify RBAC and user roles are intact by checking the platform’s Users & Access page — all roles, permission mappings, and audit log entries should match the original.
5. Verify module sources and registries
-
Restore ModuleSource resources that point to Git repositories, OCI registries, or local module directories:
kubectl apply -f /path/to/modulesources-backup.yaml -
Verify registry connectivity and pull secret validity:
kubectl -n gameplane-system get secrets | grep -E 'registry|pull-secret' kubectl -n gameplane-system describe secret <name> -
Test module availability by navigating to the Modules page and checking that all sources are reachable and module counts match the original.
6. Restore worlds into isolated servers (optional staging step)
-
Create a test GameServer and restore a snapshot of a world to it using the dashboard:
- Go to Backups and select an existing backup.
- Click Restore and choose the test server as the target.
- Monitor the Restore resource’s phase (Pending → Suspending → Running → Resuming → Succeeded).
Or use the REST API:
# Forward the API port to your local machine first (if not exposed) kubectl -n gameplane-system port-forward svc/gameplane-api 8080:80 & # Create a Restore resource via the API curl -X POST http://localhost:8080/restores \ -H "Authorization: Bearer $TOKEN" \ -H "Content-Type: application/json" \ -d '{ "backupRef": {"name": "world-backup-1"}, "serverRef": {"name": "test-server"} }' -
Verify the restored world is playable and data is intact before proceeding to production servers.
7. Restore or recreate production GameServers
-
Decide per-server: restore from a backup, or re-provision and let players rejoin.
- Restore from backup: fastest for world data, but assumes the backup is recent and valid.
- Recreate: safest for platform bugs or data corruption; ensures all configuration is fresh.
-
Restore into production servers:
# Re-apply all backed-up GameServers (they will start without world data first) kubectl apply -f /path/to/gameservers-backup.yaml # For each world that has a corresponding backup, trigger a restore via the dashboard: # Backups → select backup → Restore → choose target server → confirmAlternatively, use the REST API to trigger restores programmatically:
kubectl -n gameplane-system port-forward svc/gameplane-api 8080:80 & # For each world backup curl -X POST http://localhost:8080/restores \ -H "Authorization: Bearer $TOKEN" \ -H "Content-Type: application/json" \ -d '{ "backupRef": {"name": "world-backup-1"}, "serverRef": {"name": "server-1"} }' -
Reconnect endpoints and re-enable player access once servers are ready.
Validate and rehearse
Compare the rebuilt environment to the recovery manifest and run scheduled isolated drills.
Validation checklist
- Platform state: Users, roles, audit logs, settings, and module sources match the original.
- Cluster topology: Multi-cluster Clusters are registered and can reach remote kubeconfigs.
- Storage and backups: StorageClass is correct, backup destinations are reachable, and a new backup can be created.
- Networking and TLS: Ingress is reachable at the original hostname, TLS certs are valid, and DNS resolves.
- Monitoring and audit: Webhook/syslog/S3 audit delivery is active, metrics are flowing to Prometheus, and alerts are firing as expected.
Disaster recovery drills
- Quarterly or twice yearly: run a full recovery drill on a staging cluster using a dated backup.
- Document the actual RPO and RTO: measure how long each phase takes (database restore, CRD restoration, GameServer setup, world restore) and record any gaps in the recovery manifest.
- Rotate recovery credentials as part of the drill; verify that old backups can still be restored with current secrets.
RECOVERY ORDER
Related guides
- Backups & Recovery — create and restore game world snapshots
- Backup Destinations — configure S3, GCS, SFTP, and local restic repositories
- Credential Rotation — rotate certificates, keys, and platform credentials
- Database Configuration & Lifecycle — configure SQLite or PostgreSQL for the API
- High Availability — replicate the control plane and database for production resilience