Observability and Evidence Architecture
Why this chapter matters
Observability is how infrastructure tells the truth about its own behaviour. Metrics, traces, logs, and evidence are useful only when their limits, privacy, retention, and interpretation remain visible to the people relying on them.
Continue to Identity and Access Architecture to connect evidence with accountable actors.
Defines how logs, metrics, traces, decisions, policy evaluations, and integrity events remain useful evidence.
- INFRA5-R001: Consequential telemetry SHALL identify source, time, scope, transformation, integrity state, and limitations.
- INFRA5-R002: Observability SHALL preserve material failures, dissent, missing data, and uncertainty rather than presenting a clean narrative.
- INFRA5-R003: Monitoring access SHALL not grant operational command or permission to alter the observed system.
- INFRA5-R004: Retention and disclosure SHALL follow purpose limitation, minimisation, legal hold, and access controls.
- INFRA5-R005: Telemetry design SHALL state coverage, sampling, clock basis, ordering, redaction, integrity, retention, access, and blind spots.
- INFRA5-R006: Monitoring SHALL distinguish absence of evidence from evidence of absence and SHALL preserve material gaps in the record.
- INFRA5-R007: Consequential alerts SHALL identify thresholds, review owner, escalation path, affected scope, and the condition for withdrawal.
- INFRA5-R008: An observability design SHALL NOT collect live telemetry or grant monitoring access operational command; it remains conceptual evidence architecture.
This Draft excludes live telemetry collection.
Evidence design
Observability shall define event classes, clock and ordering assumptions, sampling limits, redaction rules, integrity checks, retention, and independent review. A metric is a signal with a known scope, not a complete account of system or human impact.
Failure cases
Missing events, clock skew, selective sampling, forged telemetry, privacy over-collection, and monitoring blind spots shall remain visible. A monitoring gap shall reduce confidence in a conclusion rather than be reported as no incident.
Operating model and evidence
Observability design classifies events, metrics, traces, policy evaluations, decisions, and integrity signals by purpose and consequence. It records clock assumptions, sampling, aggregation, redaction, retention, access, and the limits of what each signal can establish. A clean dashboard is not a complete account when collection is partial or affected people are absent from the signal.
Reviewers compare telemetry with source events, qualitative evidence, policy records, and affected-person impact. They record missing events, clock skew, selective sampling, forged signals, alert fatigue, and privacy risk. An alert is a review input with an owner and expiry, not an operational command.
Interpretation cases
- Conforming: Coverage, timing, sampling, integrity, retention, blind spots, and review triggers are explicit.
- Prohibited: Monitoring access becomes command authority.
- Boundary: A telemetry gap narrows the conclusion rather than proving no incident.
- Failure: Clock or integrity uncertainty preserves raw evidence and pauses consequential claims.
- Loophole: Selective sampling creates a falsely clean narrative.
- Misuse: Observability data is reused for unrelated surveillance or identity exposure.
- Care-control: Monitoring supports safety while minimising collection and preserving human review.
Design evidence
Observability review should state coverage, sampling, clock basis, redaction, integrity checks, retention, access, alert thresholds, and known blind spots. Metrics require interpretation and should be compared with qualitative evidence and affected-person impact where consequence is material.