RUNBOOKS
Troubleshooting index
Triage failed starts, storage and network faults, login problems, backup failures, and module synchronization with one evidence-first workflow.
Troubleshootingv0.220 MIN
Capture evidence before making changes
Capture the resource UID, cluster, namespace, time range, events, conditions, and relevant logs before changing state.
Build the evidence bundle
Start with desired spec and observed status, then correlate events, workload state, component logs, audit records, and recent changes.
Record timestampsRecord exact timestamps and timezone for every exported source.
Separate logsSeparate Gameplane API/operator logs from game-container failures.
Preserve first errorPreserve the first error; later retries often hide the original cause.
Common failure families
Classify the symptom before choosing a runbook: Pending, image pull, crash loop, probe, volume, service, auth, backup, or source sync.
PendingUsually points to capacity, placement, PVC, or scheduling policy.
No console outputCan mean the container never started or the game stream isn't ready.
Restore and module failuresRequire destination/source credentials and controller logs.
Escalate with context
When local checks are exhausted, share a minimal reproducible case without secrets, tokens, world data, or player information.
TRIAGE ORDER
01 01 Spec → status/conditions → events → workload and PVC/service state
02 02 API/operator logs → container logs → game logs → audit history
03 03 Recent change → safe rollback → sanitized support bundle