Failure record
What has gone wrong, written down
Any operating model can be made to look good on a good day. This page is the bad days. For each one: how the defect got in, how long it ran before anybody noticed, how many units it touched, how many were reversed, and what specifically changed so it cannot happen the same way twice.
Health
What has gone wrong and what is still broken
Incidents by state, with the units of work affected and how many have been put right.
Incidents on record
16
Still open
10
actions outstanding or monitoring
Mean time to detect
3d 3h
from defect to first alarm
Mean time to contain
1d 4h
from alarm to agents paused
Units touched / reversed
31,445 / 21,438
68% recovered without a human rerun
Demand Qualification: prompt regression
Demand Qualification · lead Director, Demand Development
A change intended to improve one behavior degraded another that was not tested.
Detect
1d 5h
Contain
7h 25m
Restore
1d 23h
Root cause
A model version change altered how the agent read a free-text field. Output stayed syntactically valid and semantically wrong.
How it was detected
Found by the challenger agent during a routine dissent pass after 30 hours.
Containment
The lane was held within 16 hours. 3 agents were paused and their work re-routed to a human queue at 8 times the unit cost.
Rollback
55 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
407 units touched · 330 reversed automatically · 81% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Drift detection needs a population-level check, not a per-decision check.
- Cost of the manual fallback was 26 times the agentic unit cost. That number belongs in the business case.
Campaign Orchestration: control gap
Campaign Orchestration · lead Director, Campaign Operations
A path existed through the process with no gate on it. Nobody noticed until volume found it.
Detect
3d 16h
Contain
1d 14h
Restore
5d 13h
Root cause
Two agents each held part of a decision that should have been segregated. Neither breached its own envelope.
How it was detected
Found by a downstream reconciliation that would not tie after 88 hours.
Containment
The lane was held within 36 hours. 3 agents were paused and their work re-routed to a human queue at 17 times the unit cost.
Rollback
78 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
16 units touched · 14 reversed automatically · 88% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Every lane needs a reversal path that is exercised, not just documented.
- Cost of the manual fallback was 8 times the agentic unit cost. That number belongs in the business case.
Revenue Analytics: silent drift
Revenue Analytics · lead VP Revenue Analytics
Output quality degraded without any alarm firing. Found by sampling, not by monitoring.
Detect
2d 5h
Contain
1d 4h
Restore
2d 19h
Root cause
An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.
How it was detected
Found by an external party raising a query after 54 hours.
Containment
The lane was held within 14 hours. 2 agents were paused and their work re-routed to a human queue at 22 times the unit cost.
Rollback
83 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
2,493 units touched · 1,816 reversed automatically · 73% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Segregation of duties has to be tested across agents, not within one agent.
- Cost of the manual fallback was 10 times the agentic unit cost. That number belongs in the business case.
Compliance & Ethics: silent drift
Compliance & Ethics · lead Chief Compliance Officer
Output quality degraded without any alarm firing. Found by sampling, not by monitoring.
Detect
1d 7h
Contain
14h 4m
Restore
2d 22h
Root cause
An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.
How it was detected
Found by a downstream reconciliation that would not tie after 31 hours.
Containment
The lane was held within 18 hours. 2 agents were paused and their work re-routed to a human queue at 11 times the unit cost.
Rollback
80 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
2,754 units touched · 2,413 reversed automatically · 88% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Every lane needs a reversal path that is exercised, not just documented.
- Cost of the manual fallback was 22 times the agentic unit cost. That number belongs in the business case.
Spend Analytics: silent drift
Spend Analytics · lead Director, Procurement Analytics
Output quality degraded without any alarm firing. Found by sampling, not by monitoring.
Detect
2d 23h
Contain
12h 20m
Restore
1d 6h
Root cause
An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.
How it was detected
Found by a human who noticed the pattern in a weekly review after 72 hours.
Containment
The lane was held within 24 hours. 1 agents were paused and their work re-routed to a human queue at 18 times the unit cost.
Rollback
82 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
1,147 units touched · 742 reversed automatically · 65% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Every lane needs a reversal path that is exercised, not just documented.
- Cost of the manual fallback was 20 times the agentic unit cost. That number belongs in the business case.
Resilience & Risk: cascade
Resilience & Risk · lead Director, Supply Resilience
One tower defect propagated to every function feeding it before containment.
Detect
4d 16h
Contain
1d 3h
Restore
2d 8h
Root cause
A held unit blocked a downstream lane in another function. The queue built for two days before anyone owned it.
How it was detected
Found by a downstream reconciliation that would not tie after 112 hours.
Containment
The lane was held within 60 hours. 4 agents were paused and their work re-routed to a human queue at 15 times the unit cost.
Rollback
54 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
362 units touched · 175 reversed automatically · 48% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Drift detection needs a population-level check, not a per-decision check.
- Cost of the manual fallback was 24 times the agentic unit cost. That number belongs in the business case.
Engineering Analytics: policy version lag
Engineering Analytics · lead Director, Engineering Analytics
A rule changed in the register and agents kept enforcing the previous version.
Detect
1d 5h
Contain
17h 4m
Restore
18h 11m
Root cause
A jurisdiction changed a rule. The policy agent held the previous version for eleven days before the change was ingested.
How it was detected
Found by a downstream reconciliation that would not tie after 29 hours.
Containment
The lane was held within 7 hours. 1 agents were paused and their work re-routed to a human queue at 9 times the unit cost.
Rollback
96 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
3,720 units touched · 3,147 reversed automatically · 85% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Segregation of duties has to be tested across agents, not within one agent.
- Cost of the manual fallback was 22 times the agentic unit cost. That number belongs in the business case.
IP Creation & Capture: silent drift
IP Creation & Capture · lead Director, IP Capture
Output quality degraded without any alarm firing. Found by sampling, not by monitoring.
Detect
3d 10h
Contain
1d 23h
Restore
5d 5h
Root cause
An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.
How it was detected
Found by an external party raising a query after 82 hours.
Containment
The lane was held within 24 hours. 1 agents were paused and their work re-routed to a human queue at 5 times the unit cost.
Rollback
66 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
3,328 units touched · 2,769 reversed automatically · 83% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Every lane needs a reversal path that is exercised, not just documented.
- Cost of the manual fallback was 27 times the agentic unit cost. That number belongs in the business case.
Tax: upstream data defect
Tax · lead VP Tax
The agent behaved correctly on data that was already wrong before it arrived.
Detect
4d 9h
Contain
21h 54m
Restore
8d 9h
Root cause
A master data change propagated a wrong attribute into the lane. The agents applied the rule correctly to bad input.
How it was detected
Found by a downstream reconciliation that would not tie after 106 hours.
Containment
The lane was held within 44 hours. 2 agents were paused and their work re-routed to a human queue at 9 times the unit cost.
Rollback
80 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
372 units touched · 229 reversed automatically · 62% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Every lane needs a reversal path that is exercised, not just documented.
- Cost of the manual fallback was 12 times the agentic unit cost. That number belongs in the business case.
Culture & Engagement: prompt regression
Culture & Engagement · lead Director, Culture & Engagement
A change intended to improve one behavior degraded another that was not tested.
Detect
1d 13h
Contain
4h 29m
Restore
2d 14h
Root cause
A model version change altered how the agent read a free-text field. Output stayed syntactically valid and semantically wrong.
How it was detected
Found by a downstream reconciliation that would not tie after 38 hours.
Containment
The lane was held within 4 hours. 4 agents were paused and their work re-routed to a human queue at 17 times the unit cost.
Rollback
84 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
3,210 units touched · 2,109 reversed automatically · 66% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Drift detection needs a population-level check, not a per-decision check.
- Cost of the manual fallback was 24 times the agentic unit cost. That number belongs in the business case.
Service Desk: control gap
Service Desk · lead Director, Service Desk
A path existed through the process with no gate on it. Nobody noticed until volume found it.
Detect
3d 8h
Contain
17h 35m
Restore
4d 20h
Root cause
Two agents each held part of a decision that should have been segregated. Neither breached its own envelope.
How it was detected
Found by a downstream reconciliation that would not tie after 80 hours.
Containment
The lane was held within 35 hours. 1 agents were paused and their work re-routed to a human queue at 16 times the unit cost.
Rollback
93 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
1,630 units touched · 877 reversed automatically · 54% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Segregation of duties has to be tested across agents, not within one agent.
- Cost of the manual fallback was 4 times the agentic unit cost. That number belongs in the business case.
Real Estate Portfolio: policy version lag
Real Estate Portfolio · lead VP Corporate Real Estate
A rule changed in the register and agents kept enforcing the previous version.
Detect
2d 9h
Contain
1d 2h
Restore
3d 16h
Root cause
A jurisdiction changed a rule. The policy agent held the previous version for eleven days before the change was ingested.
How it was detected
Found by the drift monitor once the sample size finally cleared the threshold after 57 hours.
Containment
The lane was held within 25 hours. 2 agents were paused and their work re-routed to a human queue at 21 times the unit cost.
Rollback
79 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
4,094 units touched · 2,576 reversed automatically · 63% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Drift detection needs a population-level check, not a per-decision check.
- Cost of the manual fallback was 10 times the agentic unit cost. That number belongs in the business case.
Service Analytics: cascade
Service Analytics · lead Head of Service Analytics
One tower defect propagated to every function feeding it before containment.
Detect
5d 7h
Contain
2d 22h
Restore
3d 12h
Root cause
A held unit blocked a downstream lane in another function. The queue built for two days before anyone owned it.
How it was detected
Found by the drift monitor once the sample size finally cleared the threshold after 127 hours.
Containment
The lane was held within 76 hours. 4 agents were paused and their work re-routed to a human queue at 22 times the unit cost.
Rollback
95 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
2,904 units touched · 1,580 reversed automatically · 54% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Policy version ingestion needs its own SLA with a named owner.
- Cost of the manual fallback was 18 times the agentic unit cost. That number belongs in the business case.
Service Catalog & Intake: control gap
Service Catalog & Intake · lead Head of Service Catalog
A path existed through the process with no gate on it. Nobody noticed until volume found it.
Detect
5d 21h
Contain
1d 21h
Restore
6d 5h
Root cause
Two agents each held part of a decision that should have been segregated. Neither breached its own envelope.
How it was detected
Found by a human who noticed the pattern in a weekly review after 141 hours.
Containment
The lane was held within 69 hours. 2 agents were paused and their work re-routed to a human queue at 22 times the unit cost.
Rollback
73 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
554 units touched · 477 reversed automatically · 86% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Segregation of duties has to be tested across agents, not within one agent.
- Cost of the manual fallback was 6 times the agentic unit cost. That number belongs in the business case.
R2R Control Tower: cascade
R2R Control Tower · lead Tower Lead, Record-to-Report
One tower defect propagated to every function feeding it before containment.
Detect
3d 8h
Contain
1d 22h
Restore
3d 13h
Root cause
A held unit blocked a downstream lane in another function. The queue built for two days before anyone owned it.
How it was detected
Found by an external party raising a query after 80 hours.
Containment
The lane was held within 18 hours. 1 agents were paused and their work re-routed to a human queue at 20 times the unit cost.
Rollback
79 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
3,464 units touched · 1,575 reversed automatically · 45% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Segregation of duties has to be tested across agents, not within one agent.
- Cost of the manual fallback was 25 times the agentic unit cost. That number belongs in the business case.
Master Data Management: silent drift
Master Data Management · lead Head of Master Data
Output quality degraded without any alarm firing. Found by sampling, not by monitoring.
Detect
3d 8h
Contain
1d 4h
Restore
1d 9h
Root cause
An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.
How it was detected
Found by the drift monitor once the sample size finally cleared the threshold after 80 hours.
Containment
The lane was held within 20 hours. 4 agents were paused and their work re-routed to a human queue at 16 times the unit cost.
Rollback
67 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.
Agents paused during containment
990 units touched · 609 reversed automatically · 62% recovered
What changed afterwards
- Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
- Policy version ingestion needs its own SLA with a named owner.
- Cost of the manual fallback was 14 times the agentic unit cost. That number belongs in the business case.
Actions
What is waiting on a person
An incident is closed when the work it damaged has been put right, not when the system came back.
Reverse a unit of work an incident touched
Units affected that have not been put back.
Review a slow restore
Took longer than four hours to bring back.
Finish the actions from a closed incident
Service restored, remediation not done.
Decide whether an incident can be closed
Stable but still being watched.
Operations
What this desk is allowed to start
A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.
Trigger and bound
This desk can start an incident record, attach affected units to it, and log a containment or restore timestamp. It cannot reverse a transaction, restore a system, or waive a remediation action — the record follows the recovery, it does not perform it.
Live observability
What the record shows right now
Incident state. An incident with actions open is one where the service is back but the damage is not fully repaired.
Current distribution
16 incidents
Is policy and strategy coming to fruition
Whether the written intent is holding here
21,438 of 31,445 affected units have been reversed.
Not holding on the record
The written commitment is that no incident closes with unreversed work behind it. Against 16 incidents, 31,445 units of work were touched and 21,438 have been put right, leaving 10,007 outstanding. 5 incidents carry open remediation actions. Detection, containment and restore times on this page are as recorded by the responding team; none has been reviewed independently.