Cluster and storage operations
Read Kubernetes health in Gameplane, plan capacity, and choose storage behavior that supports expansion, placement, and disaster recovery.
Node status, StorageClass capabilities, and capacity metrics come from the Kubernetes API. Gameplane surfaces observability; your cluster configuration, hardware, and storage drivers determine what is possible.
Cluster preflight and health
Confirm supported Kubernetes versions, ready nodes, allocatable capacity, required controllers, and namespace access.
- Check node readiness and pressure before blaming a server workload. Gameplane surfaces node conditions (
NotReady,MemoryPressure,DiskPressure,PIDPressure) in the Cluster Operations dashboard; resolve node-level issues first. If a GameServer pod fails to schedule, check node allocatable resources (CPU, memory, ephemeral storage) viakubectl describe node. - Reserve headroom for the operator, agents, ingress, and backup jobs. The operator and API each request 100m CPU and 64Mi memory (with limits of 500m CPU and 256Mi memory). Agent sidecars in game pods also require CPU and memory; consult the agent image documentation for current resource requests. Backup and restore jobs are bursty; reserve a node or memory ceiling for their temporary pods. A single underprovisioned node can stall the entire cluster.
- Read access to cluster health is broad by default; managing nodes is not. All three built-in roles (
admin,operator,viewer) can view the Cluster Operations screen by default (cluster:read); onlyadminandoperatorcan add nodes or mint kubeconfigs (cluster:manage). Use custom roles if you need to restrict read access further.
Kubernetes and Helm versions: Gameplane requires Kubernetes 1.28+ and Helm 3.13+. Check your cluster’s version:
kubectl version --short
helm version --short
If your cluster is older, upgrade it before installing or upgrading Gameplane. Consult your Kubernetes distribution’s documentation (k3s, EKS, GKE, AKS) for upgrade steps.
Persistent storage
Match access modes, reclaim policy, expansion, topology, and backup strategy to the game’s world data.
- Use a default StorageClass or select one explicitly during creation. Select a StorageClass explicitly per server or template via
spec.storage.storageClassName; if none is set, the cluster’s default StorageClass is used. Verify one exists and matches your needs (e.g., SSD for high-IOPS games, networked storage for multi-pod backups).
An install-time Helm default (operator.gameDataStorage.storageClassName) is coming in v0.3.0 and is not yet available in beta.8.
- Confirm
allowVolumeExpansionbefore increasing a server PVC. Game worlds grow; if you plan to increase PVC size later, ensure the StorageClass allows expansion. Check withkubectl get storageclass -o wide— theALLOWVOLUMEEXPANSIONcolumn should betrue. You can edit a GameServer’sspec.storage.sizeto trigger a resize (the Kubernetes CSI driver and underlying storage backend must support it). - Understand whether deletion retains or removes the underlying volume. StorageClass
reclaimPolicyisDelete(volume deleted on PVC deletion) orRetain(volume kept for manual recovery). For production game servers, considerRetainto preserve world data even if a GameServer CRD is deleted. Consult your storage provider’s documentation on reclaim semantics.
Backup and storage lifecycle: Backups are created on-demand or via schedule; they are stored in a backup destination (S3, Azure Blob, local restic repository, etc.), not on the PVC. Restoration creates a new GameServer with a new PVC and restores world data from the backup; the original PVC remains untouched until explicitly deleted.
Placement and multi-node fleets
Use scheduler defaults first, then add affinity, tolerations, GPU selectors, or pinning only for a clear constraint.
Start with Kubernetes’ built-in scheduler — it balances pods across nodes, respects resource requests, and schedules based on node selectors and tolerations. Only add affinity rules if you have a hard constraint (e.g., “all game pods for server X must run on the same node” or “high-memory servers must use nodes with NVMe”).
Affinity: Define pod affinity to co-locate servers on the same node (cheaper networking, lower latency) or spread them across different zones (redundancy). Affinity is a soft preference (preferredDuringSchedulingIgnoredDuringExecution) or hard requirement (requiredDuringSchedulingIgnoredDuringExecution). Hard affinity can starve other workloads if not carefully scoped.
Tolerations and taints: If a node has a taint (e.g., dedicated=game-servers:NoSchedule), game pods need a matching toleration to schedule on it. Taints are useful for reserving nodes for a specific workload type.
GPU and resource selectors: If a game template uses GPU (e.g., for rendering or simulation), add a node selector or request the GPU resource in the template’s spec.resources.limits and spec.resources.requests. Gameplane will schedule the pod on a node with available GPU capacity.
Affinity best practice: In a multi-node fleet, avoid hard affinity requirements on individual servers — let the scheduler spread pods to maximize cluster resilience. Use soft affinity (preferred) for co-location where latency matters.
CLUSTER PREFLIGHT
Related guides
- Multi-cluster Topology — register remote Kubernetes clusters and distribute workloads across them
- High Availability — replicate the control plane and database for production resilience
- Kubernetes Maintenance — plan node upgrades and cluster patches
- Database Configuration & Lifecycle — configure SQLite or PostgreSQL for the API
- Backups & Recovery — create and restore game world snapshots