Operations & Observability
Overview
Operations and observability standards require production services to expose their behaviour and support operational response and recovery.
Choose a standard below to read it in full.
Standards
Structured Logging
A service's log output is structured, traceable, and free of sensitive data or unnecessary detail.
Metrics, Monitoring & Alerting
A service's metrics remain accurate, visible, and bounded, with automated and actionable alerts.
Distributed Tracing
Trace context propagates end-to-end through accurately structured spans and deliberate sampling.
Telemetry Instrumentation
Telemetry remains consistent across services, safe to introduce, and delivered alongside the change it observes.
Observability Platform Integration
Telemetry is centralised, portable, and available for resilient operations.
Backup & Disaster Recovery
Recovery is designed and validated against defined objectives.
Runbooks
A runbook documents and validates the steps to resolve known failures and perform high-risk procedures.