gameplane / docs
PLATFORM

Kubernetes node maintenance and upgrades

Drain nodes and upgrade Kubernetes without corrupting worlds, stranding volumes, or exceeding supported disruption limits.

Cluster & Storagev0.218 MIN

Cluster maintenance involves draining nodes for patches, version upgrades, or hardware replacement. This page covers the safe sequence: inventory your workload, cordon the node, migrate or stop servers, drain gracefully, upgrade the node, then return to service.

Portability isn't automatic

Portability depends on storage topology, GPU and port constraints, affinity, and real spare capacity—not just a Ready destination node.

Prepare the window

Confirm compatibility, capacity, workload placement, storage mobility, maintenance communication, and recent backups.

Check Kubernetes versionsVerify Kubernetes, CSI, ingress, cert-manager, and rollback support.
Inventory affected serversAnnounce, save worlds, and verify backups.
Identify non-movable resourcesIdentify local volumes, GPUs, NodePorts, affinity, and non-movable pods.

Before cordoning any node, check whether your cluster can tolerate the loss:

  • List all GameServers and Pods scheduled on the target node using kubectl get pod -n gameplane-games --field-selector=spec.nodeName=<node-name>.
  • Check storage: identify PVCs mounted by those Pods (especially local storage or single-replica StatefulSets).
  • Document affinity, NodePorts, and any workload that cannot migrate (e.g., a singular StatefulSet without spread).
  • Ensure backups of affected servers are recent and healthy.
  • Announce the maintenance window to players and server operators.

Cordon, drain, and upgrade

Move one failure domain at a time using deliberate eviction behavior and provider-supported ordering.

Cordon the nodeCordon, then stop or migrate servers according to storage behavior.
Drain gracefullyDrain with explicit flags/timeouts; never force-delete state blindly.
Monitor healthObserve API, operator, agent, CSI, and ingress health during upgrade.

Cordon the node

A cordoned node will not accept new Pods but keeps running ones intact:

kubectl cordon <node-name>

Migrate or stop servers

Next, handle the GameServers running on that node. The strategy depends on your storage:

  • Servers with remote storage (NFS, EBS, cloud-managed PVC): The Pod can restart on another node. Trigger a graceful stop by setting spec.suspend: true on the GameServer (scales its StatefulSet to zero while preserving data), or let the drain process evict it.
  • Servers with local storage: Stop the Pod before draining, save its world/data to remote backup, and plan to restart it after the node returns.

Drain the node

Drain evicts all Pods from the node, respecting Pod Disruption Budgets (PDBs) and graceful shutdown:

kubectl drain <node-name> \
  --ignore-daemonsets \
  --ignore-errors \
  --delete-emptydir-data \
  --grace-period=300

Flags explained:

  • --ignore-daemonsets: Keep DaemonSet Pods (e.g., CNI, kubelet agents) running; they auto-restart on the upgraded node anyway.
  • --ignore-errors: Skip non-evictable Pods (rare; usually means something is not a standard Kubernetes object). Continue draining other Pods.
  • --delete-emptydir-data: Permit eviction of Pods using emptyDir storage (safe if data is non-persistent).
  • --grace-period=300: Allow 5 minutes for Pods to shut down cleanly. Set higher for stateful servers that need time to save.

Upgrade the node

Once drained, the node is safe to reboot or replace:

# Option A: in-place upgrade (e.g., apt/yum upgrade)
sudo apt-get update && sudo apt-get upgrade && sudo reboot

# Option B: node replacement (e.g., cloud provider instance replacement)
kubectl delete node <node-name>  # Node will rejoin after replacement/restart

Monitor health during upgrade

Watch the cluster and Gameplane components:

# API server readiness
kubectl get deployment -n gameplane-system gameplane-api

# Operator health
kubectl get deployment -n gameplane-system gameplane-operator

# CSI health (if using external storage)
kubectl get pod -n kube-system | grep csi

# Ingress readiness
kubectl get ingress -n gameplane-games

Return to service

Uncordon only after node, runtime, network, CSI, certs, and capacity are healthy, then validate every affected workflow.

Verify the upgraded node is ready:

kubectl get node <node-name>

The node should show Status: Ready and all Pods in Ready phase. Then uncordon:

kubectl uncordon <node-name>

Validate after uncordon

  • Pod spreading: Re-schedule Pods that were held during drain. The scheduler will now place new Pods on the uncordoned node.
  • PVC attachment: If using remote storage, verify all PVCs re-attached and are readable (check kubectl get pvc and pod logs).
  • Endpoints: For servers exposed via NodePort or LoadBalancer, ensure the node’s IP is in the endpoint list (kubectl get endpoints -n gameplane-games).
  • Backups: Run a post-upgrade backup validation; ensure recent snapshots are healthy.
  • Monitoring & alerts: Check dashboards for pod restart rate, latency spikes, or any new alerts.

MAINTENANCE LOOP

01   Compatibility + capacity + movable storage + backups gate
02   Cordon → save/stop/migrate → drain → upgrade → health check
03   Uncordon → validate pods, PVCs, endpoints, backups, alerts, audit