Production observability and alerting
Connect metrics, logs, events, audit records, and notifications into an actionable signal path with bounded retention.
Every alert needs an owner, duration, severity, notification route, and tested runbook—not just a threshold.
Collect the right signals
Label every signal with cluster, namespace, resource UID, and synchronized time while keeping component boundaries visible.
Metrics endpoints:
- API:
/metricson the dedicated in-cluster listener (portmetricsPort, default 9090). ServiceMonitor available whenserviceMonitors.enabled: trueinvalues.yaml. - Operator:
/metricson port 8080. Exports fleet state gauges likegameplane_gameservers{phase=...}and reconciliation counters. - Agent sidecars:
/metricson port 9090 (unauthenticated). PodMonitor available whenserviceMonitors.enabled: true.
Logs collection:
Configure your cluster’s log aggregation system (Kubernetes logging architecture) to collect:
- API, operator, and agent stdout/stderr from
kubectl logs - Game server logs from the agent’s tail stream (available via the
/ws/servers/<name>/logsWebSocket endpoint used by the dashboard’s Logs tab)
Audit trail:
Export audit events via webhook syslog (RFC 5424) using api.audit.webhook.syslogBridge or batch upload to S3 using api.audit.s3.* in values.yaml. See Audit observability for full configuration.
Dashboards, SLOs, and alerts
Cover control-plane health, reconciliation, fleet states, capacity, storage, backups, credentials, and source synchronization.
Key alerts to define:
- Reconciliation health:
GameplaneOperatorReconcileErrors(by controller),GameplaneOperatorReconcileStuck(longest-running work > 5m) - Fleet health:
GameplaneGameServerFailed(phase=Failed),GameplaneBackupFailed(any failed backup) - Capacity: API/operator memory, CPU throttling, PVC usage approaching limits
- Infrastructure: Certificate expiry, Database connectivity, Persistent volume mount failures
- Sources & modules: Module source sync failures, registry credential rotation expiry (see Credential rotation)
Built-in PrometheusRule:
Enable prometheusRules.enabled: true in values.yaml to activate the bundled alert rules. Use serviceMonitors.enabled: true to auto-create Prometheus scrape targets and scrapeNamespaceSelector to open NetworkPolicy exceptions if Prometheus runs in a separate namespace.
Retention and incident use
Bound cardinality, access, redaction, encryption, and cost; preserve evidence before restarts or rotation.
Metrics retention: Prometheus is stateful; size your Persistent Volume for your retention window (commonly 15d–90d depending on cost/compliance). High-cardinality labels (e.g., per-pod or per-request-id) increase storage; use recording rules to pre-aggregate before long-term storage.
Logs retention: Cluster log aggregators (Loki, ELK, Stackdriver) typically offer configurable retention (7d–30d by default). Game logs are available on-demand via the dashboard; configure agent log persistence (spec.logPath in templates) to disk (attached volume) or external log sink.
Audit trail: API audit events are persisted in the database (every row hash-chained per 005_audit_chain.sql migration). Export via syslog bridge or S3 for long-term archival and legal/compliance holds. Gameplane never deletes audit rows—you control archival and purge policies.
Evidence preservation: Before node maintenance, cluster upgrades, or credential rotation, export:
- Prometheus snapshots (if using Prometheus)
- Recent log chunks from your aggregator
- Audit event exports from the API (
GET /admin/audit/exportor syslog archive) - Game server snapshots (if applicable for your incident timeline)