Backup backends and repository operations
Configure destinations with correct credentials, encryption, retention, health checks, and recurring restore validation.
A Backup or BackupSchedule references a repository via a Kubernetes Secret containing the restic repository URI and password. This guide covers supported backends, credential management, operational monitoring, and restore validation drills.
Choose and configure a backend
Document supported repository type, endpoint, TLS, permissions, immutability, lifecycle, and environment separation.
Repository URI formats
Restic repository URIs follow the pattern <backend>:<backend-specific-config>. Common examples:
- S3-compatible:
s3:s3.amazonaws.com/bucket-name(AWS S3),s3:minio-host:9000/bucket-name(MinIO),s3:s3.backblazeb2.com/bucket-name(B2) - REST:
rest:https://restic-server.example.com:8000/(requires arestic rest-serverinstance) - SFTP:
sftp:user@sftp.example.com:22/path/to/repo - Filesystem:
local:/mnt/backup-pvc(for local filesystem or PVC-mounted backends)
Restic also supports cloud-native backends (Azure Blob, Google Cloud Storage, Swift) via environment variables; see the Restic documentation for backend-specific setup.
Gameplane’s Backup/BackupSchedule Secret exposes only two keys to the restic Job: repo (the repository URI) and password (the restic repository’s encryption password, mapped to RESTIC_PASSWORD) — this is unrelated to any backend access key. The operator does not inject backend-specific credentials (AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, Azure/GCS/Swift env vars) into the Job, so backends that require them need another route (e.g., node/pod identity such as IRSA) since Gameplane passes no extra environment variables to the restic container.
Operate the repository
Define initialization, locks, checks, prune, retention, concurrency, bandwidth, and credential rotation.
Repository health checks
Restic’s check command verifies repository integrity. Gameplane does not automate this; add a separate CronJob or scheduled Restore+verification job to periodically run restic check against each repository:
apiVersion: batch/v1
kind: CronJob
metadata:
name: backup-repo-health-check
namespace: gameplane-games
spec:
schedule: "0 2 * * 0" # Weekly at 2 AM UTC
jobTemplate:
spec:
template:
spec:
serviceAccountName: backup-job
containers:
- name: restic-check
image: restic/restic:0.19.1
env:
- name: RESTIC_REPOSITORY
valueFrom:
secretKeyRef:
name: backup-repo-secret
key: repo
- name: RESTIC_PASSWORD
valueFrom:
secretKeyRef:
name: backup-repo-secret
key: password
command: ["restic", "check", "--read-data-subset=10%"]
restartPolicy: OnFailure
Pruning and retention
Restic’s retention policies are defined in BackupSchedule.spec.retention. Gameplane supports restic’s keep-* vocabulary:
keepLast: 5— always keep the 5 most recent snapshots regardless of agekeepHourly: 24— keep the latest snapshot from each of the last 24 hourskeepDaily: 7— keep the latest snapshot from each day (7 daily snapshots)keepWeekly: 4— keep the latest snapshot from each week (4 weekly snapshots)keepMonthly: 12— keep the latest snapshot from each month (12 monthly snapshots)keepYearly: 3— keep the latest snapshot from each year (3 yearly snapshots)
The operator trims retention by deleting old Backup custom resources once they fall outside every configured keep-* bucket (surfaced via the RetentionTrimmed condition on the BackupSchedule). It does not run restic forget/prune against the repository itself, so repository storage is not automatically reclaimed. If you need to reclaim space, run restic forget --prune yourself on a schedule that matches your retention policy, and monitor that job’s own logs for bytes freed and errors.
Credentials and certificate rotation
Restic repository credentials (S3 access keys, SFTP passwords, API tokens) and TLS certificates should be rotated regularly (e.g., quarterly or per compliance policy). To rotate:
- Create a new Secret with updated credentials.
- Update the Backup/BackupSchedule to reference the new Secret’s name in
spec.repoRef. - Test a backup with the new credentials to verify access works.
- Delete the old Secret only after confirming all dependent backups are using the new one.
During rotation, the repository itself remains accessible to existing backups; the rotation is a Kubernetes Secret update, not a repository migration.
Prove restoreability
Restore representative worlds on schedule using the exact key and access procedure available during an emergency.
A backup repository is only considered “in service” after a restore drill has succeeded. Drills verify that:
- The documented credentials actually work.
- The restore process completes within your RTO target.
- The restored data is usable (game world loads, player data is intact).
Restore drill procedure
- Schedule a weekly or monthly restore test (separate from production restores) targeting a test GameServer or throwaway instance.
- Document the exact procedure — which credentials, which backup schedule/snapshot, how to authenticate if multi-factor auth or network restrictions apply.
- Simulate failure conditions: use an expired certificate, a rotated access key, or network-isolated access to verify your documented emergency procedure is complete and accurate.
- Record metrics: restore start, completion time, bytes restored, and any warnings or errors.
- Verify data integrity: confirm the restored world loads, players can connect, and data consistency checks pass.
Emergency-access credentials
Backups should be restorable using a separate, minimal set of credentials reserved for emergencies:
- A read-only access key (for S3) or SSH key (for SFTP) that only grants access to the backup repository.
- Credentials stored in a secure location (e.g., encrypted password vault, hardware security module) separate from day-to-day backup management.
- Periodically test that emergency credentials actually work, even if the primary access key changes.
This separation ensures that a compromised workload credential does not also compromise your recovery path.
REPOSITORY ACCEPTANCE
Related guides
- Backup Destinations — configure and test backup targets
- Backups & Recovery — on-demand and scheduled backup overview
- Production Observability — monitoring backup metrics and alerts