Operate · item 27
Platform continuity
Continuity of the platform is held separately from continuity of the functions it runs, because the two fail differently. A function fails when work stops. The platform can fail while every queue keeps moving — observability goes dark and the work carries on unwatched, or the evidence store stops writing and the work carries on unprovable. Those failures need their own inventory, their own tiers and their own drills.
Health
What survives the loss of a component
Components with a proven, partial or unproven recovery path, and the lines that cannot be reversed.
Function continuity asks what happens to invoices, hires and orders when a line goes down, and the answer is usually a manual fallback and a backlog. Those questions are answered per line on the drill record.
Platform continuity asks a different question: what happens when the machinery that runs, watches and proves the work goes down. The worst platform failures are silent. Losing observability does not stop a single transaction — it stops anyone from knowing whether the transactions were right, and the estate keeps running at full speed with nobody watching.
Tier 1 — Everything stops without it.
5 components
Tier 2 — A capability is lost and processing continues.
4 components
Tier 3 — Convenience or reporting only.
2 components
Platform drills on the calendar
7 drills scoped to the platform rather than to a function
Actions
What is waiting on a person
Continuity work is slow and nobody asks for it until the day it matters. These are the items that would hurt.
Operations
What this desk is allowed to start
A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.
Trigger and bound
This desk can start a component assessment and record a proven recovery objective against evidence from a drill. It cannot trigger a failover, change an architecture, or declare a component resilient without a rehearsal behind it.
Live observability
What the record shows right now
Component recovery status. Proven means a rehearsal met the stated objective.
Current distribution
11 components
Is policy and strategy coming to fruition
Whether the written intent is holding here
1 of 11 components have a proven recovery path. 6 remain single points of failure.
Not holding on the record
The written intent is that this platform can lose any one component and keep running. That is not yet demonstrated: 7 of 11 components are unproven, 2 have no failover mode configured and 6 are single points of failure. Separately, 93 business lines have no reversal path at all and 118 can only be partly reversed — which means for those lines a recovery restores the system without restoring the work.
Drill results are measured against a modeled estate on a rehearsal calendar of our own making. Recovery times here are real measurements of a simulation, not evidence that a customer's production estate recovers in the same time.
6 of the 11 platform components are single points of failure, 4 of them tier 1. 8 have a failover path that has never been exercised and 2 have no failover path at all. Those are stated as findings, not as a maturity score.
Recovery point targets on this page are targets. Only recovery time has been measured, and only on the components with a proven figure — everywhere else the field reads never proven rather than showing the target as though it were an outcome.
Vendor lock is recorded on 5 components and 2 of those have no exit plan written down. Portability is a continuity question, not a procurement one, which is why it sits here.