LensReading which lens this session carries.

Failure record

What has gone wrong, written down

Any operating model can be made to look good on a good day. This page is the bad days. For each one: how the defect got in, how long it ran before anybody noticed, how many units it touched, how many were reversed, and what specifically changed so it cannot happen the same way twice.

Health

What has gone wrong and what is still broken

Incidents by state, with the units of work affected and how many have been put right.

Incidents on record

16

Still open

10

actions outstanding or monitoring

Mean time to detect

3d 3h

from defect to first alarm

Mean time to contain

1d 4h

from alarm to agents paused

Units touched / reversed

31,445 / 21,438

68% recovered without a human rerun

INC-SL010prompt regressionMonitoringseverity medium

Demand Qualification: prompt regression

Demand Qualification · lead Director, Demand Development

A change intended to improve one behavior degraded another that was not tested.

Detect

1d 5h

Contain

7h 25m

Restore

1d 23h

Root cause

A model version change altered how the agent read a free-text field. Output stayed syntactically valid and semantically wrong.

How it was detected

Found by the challenger agent during a routine dissent pass after 30 hours.

Containment

The lane was held within 16 hours. 3 agents were paused and their work re-routed to a human queue at 8 times the unit cost.

Rollback

55 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Lead Enrichment AgentFit Scoring AgentIntent Signal Agent

407 units touched · 330 reversed automatically · 81% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Drift detection needs a population-level check, not a per-decision check.
  • Cost of the manual fallback was 26 times the agentic unit cost. That number belongs in the business case.
detected Jan 24, 2026line map related escalations
INC-MK011control gapActions openseverity low

Campaign Orchestration: control gap

Campaign Orchestration · lead Director, Campaign Operations

A path existed through the process with no gate on it. Nobody noticed until volume found it.

Detect

3d 16h

Contain

1d 14h

Restore

5d 13h

Root cause

Two agents each held part of a decision that should have been segregated. Neither breached its own envelope.

How it was detected

Found by a downstream reconciliation that would not tie after 88 hours.

Containment

The lane was held within 36 hours. 3 agents were paused and their work re-routed to a human queue at 17 times the unit cost.

Rollback

78 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Calendar Coordination AgentCross-channel Sequencing AgentAudience Suppression Agent

16 units touched · 14 reversed automatically · 88% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Every lane needs a reversal path that is exercised, not just documented.
  • Cost of the manual fallback was 8 times the agentic unit cost. That number belongs in the business case.
detected Jun 29, 2026line map related escalations
INC-RO012silent driftClosedseverity high

Revenue Analytics: silent drift

Revenue Analytics · lead VP Revenue Analytics

Output quality degraded without any alarm firing. Found by sampling, not by monitoring.

Detect

2d 5h

Contain

1d 4h

Restore

2d 19h

Root cause

An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.

How it was detected

Found by an external party raising a query after 54 hours.

Containment

The lane was held within 14 hours. 2 agents were paused and their work re-routed to a human queue at 22 times the unit cost.

Rollback

83 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Cohort Analysis Agent

2,493 units touched · 1,816 reversed automatically · 73% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Segregation of duties has to be tested across agents, not within one agent.
  • Cost of the manual fallback was 10 times the agentic unit cost. That number belongs in the business case.
detected Feb 7, 2026line map related escalations
INC-LG013silent driftActions openseverity high

Compliance & Ethics: silent drift

Compliance & Ethics · lead Chief Compliance Officer

Output quality degraded without any alarm firing. Found by sampling, not by monitoring.

Detect

1d 7h

Contain

14h 4m

Restore

2d 22h

Root cause

An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.

How it was detected

Found by a downstream reconciliation that would not tie after 31 hours.

Containment

The lane was held within 18 hours. 2 agents were paused and their work re-routed to a human queue at 11 times the unit cost.

Rollback

80 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Policy Attestation Agent

2,754 units touched · 2,413 reversed automatically · 88% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Every lane needs a reversal path that is exercised, not just documented.
  • Cost of the manual fallback was 22 times the agentic unit cost. That number belongs in the business case.
detected Jul 6, 2026line map related escalations
INC-PR014silent driftMonitoringseverity medium

Spend Analytics: silent drift

Spend Analytics · lead Director, Procurement Analytics

Output quality degraded without any alarm firing. Found by sampling, not by monitoring.

Detect

2d 23h

Contain

12h 20m

Restore

1d 6h

Root cause

An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.

How it was detected

Found by a human who noticed the pattern in a weekly review after 72 hours.

Containment

The lane was held within 24 hours. 1 agents were paused and their work re-routed to a human queue at 18 times the unit cost.

Rollback

82 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Spend Classification AgentTail Spend Agent

1,147 units touched · 742 reversed automatically · 65% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Every lane needs a reversal path that is exercised, not just documented.
  • Cost of the manual fallback was 20 times the agentic unit cost. That number belongs in the business case.
detected Apr 30, 2026line map related escalations
INC-SC015cascadeMonitoringseverity medium

Resilience & Risk: cascade

Resilience & Risk · lead Director, Supply Resilience

One tower defect propagated to every function feeding it before containment.

Detect

4d 16h

Contain

1d 3h

Restore

2d 8h

Root cause

A held unit blocked a downstream lane in another function. The queue built for two days before anyone owned it.

How it was detected

Found by a downstream reconciliation that would not tie after 112 hours.

Containment

The lane was held within 60 hours. 4 agents were paused and their work re-routed to a human queue at 15 times the unit cost.

Rollback

54 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Disruption Sensing AgentMulti-tier Mapping Agent

362 units touched · 175 reversed automatically · 48% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Drift detection needs a population-level check, not a per-decision check.
  • Cost of the manual fallback was 24 times the agentic unit cost. That number belongs in the business case.
detected Apr 17, 2026line map related escalations
INC-EN016policy version lagActions openseverity high

Engineering Analytics: policy version lag

Engineering Analytics · lead Director, Engineering Analytics

A rule changed in the register and agents kept enforcing the previous version.

Detect

1d 5h

Contain

17h 4m

Restore

18h 11m

Root cause

A jurisdiction changed a rule. The policy agent held the previous version for eleven days before the change was ingested.

How it was detected

Found by a downstream reconciliation that would not tie after 29 hours.

Containment

The lane was held within 7 hours. 1 agents were paused and their work re-routed to a human queue at 9 times the unit cost.

Rollback

96 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Delivery Metrics AgentCycle Time Agent

3,720 units touched · 3,147 reversed automatically · 85% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Segregation of duties has to be tested across agents, not within one agent.
  • Cost of the manual fallback was 22 times the agentic unit cost. That number belongs in the business case.
detected Nov 7, 2025line map related escalations
INC-RD017silent driftActions openseverity high

IP Creation & Capture: silent drift

IP Creation & Capture · lead Director, IP Capture

Output quality degraded without any alarm firing. Found by sampling, not by monitoring.

Detect

3d 10h

Contain

1d 23h

Restore

5d 5h

Root cause

An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.

How it was detected

Found by an external party raising a query after 82 hours.

Containment

The lane was held within 24 hours. 1 agents were paused and their work re-routed to a human queue at 5 times the unit cost.

Rollback

66 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Invention Detection Agent

3,328 units touched · 2,769 reversed automatically · 83% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Every lane needs a reversal path that is exercised, not just documented.
  • Cost of the manual fallback was 27 times the agentic unit cost. That number belongs in the business case.
detected Oct 4, 2025line map related escalations
INC-FI018upstream data defectMonitoringseverity medium

Tax: upstream data defect

Tax · lead VP Tax

The agent behaved correctly on data that was already wrong before it arrived.

Detect

4d 9h

Contain

21h 54m

Restore

8d 9h

Root cause

A master data change propagated a wrong attribute into the lane. The agents applied the rule correctly to bad input.

How it was detected

Found by a downstream reconciliation that would not tie after 106 hours.

Containment

The lane was held within 44 hours. 2 agents were paused and their work re-routed to a human queue at 9 times the unit cost.

Rollback

80 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Provision AgentCompliance Calendar Agent

372 units touched · 229 reversed automatically · 62% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Every lane needs a reversal path that is exercised, not just documented.
  • Cost of the manual fallback was 12 times the agentic unit cost. That number belongs in the business case.
detected May 7, 2026line map related escalations
INC-HR019prompt regressionClosedseverity high

Culture & Engagement: prompt regression

Culture & Engagement · lead Director, Culture & Engagement

A change intended to improve one behavior degraded another that was not tested.

Detect

1d 13h

Contain

4h 29m

Restore

2d 14h

Root cause

A model version change altered how the agent read a free-text field. Output stayed syntactically valid and semantically wrong.

How it was detected

Found by a downstream reconciliation that would not tie after 38 hours.

Containment

The lane was held within 4 hours. 4 agents were paused and their work re-routed to a human queue at 17 times the unit cost.

Rollback

84 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Sentiment Sensing AgentSurvey Design & Run Agent

3,210 units touched · 2,109 reversed automatically · 66% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Drift detection needs a population-level check, not a per-decision check.
  • Cost of the manual fallback was 24 times the agentic unit cost. That number belongs in the business case.
detected Feb 15, 2026line map related escalations
INC-IT020control gapActions openseverity high

Service Desk: control gap

Service Desk · lead Director, Service Desk

A path existed through the process with no gate on it. Nobody noticed until volume found it.

Detect

3d 8h

Contain

17h 35m

Restore

4d 20h

Root cause

Two agents each held part of a decision that should have been segregated. Neither breached its own envelope.

How it was detected

Found by a downstream reconciliation that would not tie after 80 hours.

Containment

The lane was held within 35 hours. 1 agents were paused and their work re-routed to a human queue at 16 times the unit cost.

Rollback

93 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Intake & Classification AgentTier-1 Resolution AgentKnowledge Retrieval Agent

1,630 units touched · 877 reversed automatically · 54% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Segregation of duties has to be tested across agents, not within one agent.
  • Cost of the manual fallback was 4 times the agentic unit cost. That number belongs in the business case.
detected May 14, 2026line map related escalations
INC-AD021policy version lagClosedseverity high

Real Estate Portfolio: policy version lag

Real Estate Portfolio · lead VP Corporate Real Estate

A rule changed in the register and agents kept enforcing the previous version.

Detect

2d 9h

Contain

1d 2h

Restore

3d 16h

Root cause

A jurisdiction changed a rule. The policy agent held the previous version for eleven days before the change was ingested.

How it was detected

Found by the drift monitor once the sample size finally cleared the threshold after 57 hours.

Containment

The lane was held within 25 hours. 2 agents were paused and their work re-routed to a human queue at 21 times the unit cost.

Rollback

79 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Lease Abstraction Agent

4,094 units touched · 2,576 reversed automatically · 63% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Drift detection needs a population-level check, not a per-decision check.
  • Cost of the manual fallback was 10 times the agentic unit cost. That number belongs in the business case.
detected May 27, 2026line map related escalations
INC-CS022cascadeClosedseverity high

Service Analytics: cascade

Service Analytics · lead Head of Service Analytics

One tower defect propagated to every function feeding it before containment.

Detect

5d 7h

Contain

2d 22h

Restore

3d 12h

Root cause

A held unit blocked a downstream lane in another function. The queue built for two days before anyone owned it.

How it was detected

Found by the drift monitor once the sample size finally cleared the threshold after 127 hours.

Containment

The lane was held within 76 hours. 4 agents were paused and their work re-routed to a human queue at 22 times the unit cost.

Rollback

95 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Metric Assembly AgentCost-to-Serve AgentContact Driver Agent

2,904 units touched · 1,580 reversed automatically · 54% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Policy version ingestion needs its own SLA with a named owner.
  • Cost of the manual fallback was 18 times the agentic unit cost. That number belongs in the business case.
detected May 29, 2026line map related escalations
INC-GB023control gapClosedseverity medium

Service Catalog & Intake: control gap

Service Catalog & Intake · lead Head of Service Catalog

A path existed through the process with no gate on it. Nobody noticed until volume found it.

Detect

5d 21h

Contain

1d 21h

Restore

6d 5h

Root cause

Two agents each held part of a decision that should have been segregated. Neither breached its own envelope.

How it was detected

Found by a human who noticed the pattern in a weekly review after 141 hours.

Containment

The lane was held within 69 hours. 2 agents were paused and their work re-routed to a human queue at 22 times the unit cost.

Rollback

73 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Catalog Definition Agent

554 units touched · 477 reversed automatically · 86% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Segregation of duties has to be tested across agents, not within one agent.
  • Cost of the manual fallback was 6 times the agentic unit cost. That number belongs in the business case.
detected Nov 21, 2025line map related escalations
INC-GB024cascadeMonitoringseverity high

R2R Control Tower: cascade

R2R Control Tower · lead Tower Lead, Record-to-Report

One tower defect propagated to every function feeding it before containment.

Detect

3d 8h

Contain

1d 22h

Restore

3d 13h

Root cause

A held unit blocked a downstream lane in another function. The queue built for two days before anyone owned it.

How it was detected

Found by an external party raising a query after 80 hours.

Containment

The lane was held within 18 hours. 1 agents were paused and their work re-routed to a human queue at 20 times the unit cost.

Rollback

79 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Journal Entry AgentReconciliation AgentIntercompany Agent

3,464 units touched · 1,575 reversed automatically · 45% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Segregation of duties has to be tested across agents, not within one agent.
  • Cost of the manual fallback was 25 times the agentic unit cost. That number belongs in the business case.
detected Apr 6, 2026line map related escalations
INC-GB025silent driftClosedseverity medium

Master Data Management: silent drift

Master Data Management · lead Head of Master Data

Output quality degraded without any alarm firing. Found by sampling, not by monitoring.

Detect

3d 8h

Contain

1d 4h

Restore

1d 9h

Root cause

An agent kept clearing units at high confidence while the underlying distribution moved. Nothing alerted because every individual decision looked normal.

How it was detected

Found by the drift monitor once the sample size finally cleared the threshold after 80 hours.

Containment

The lane was held within 20 hours. 4 agents were paused and their work re-routed to a human queue at 16 times the unit cost.

Rollback

67 percent of affected units were reversed automatically from the evidence log. The remainder needed a human because the downstream system had already settled.

Agents paused during containment

Data Quality AgentGolden Record AgentDuplicate Resolution Agent

990 units touched · 609 reversed automatically · 62% recovered

What changed afterwards

  • Confidence is not accuracy. High confidence on a shifted distribution is the failure mode, not the safeguard.
  • Policy version ingestion needs its own SLA with a named owner.
  • Cost of the manual fallback was 14 times the agentic unit cost. That number belongs in the business case.
detected Jul 6, 2026line map related escalations

Actions

What is waiting on a person

An incident is closed when the work it damaged has been put right, not when the system came back.

Operations

What this desk is allowed to start

A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.

Trigger and bound

This desk can start an incident record, attach affected units to it, and log a containment or restore timestamp. It cannot reverse a transaction, restore a system, or waive a remediation action — the record follows the recovery, it does not perform it.

Live observability

What the record shows right now

Incident state. An incident with actions open is one where the service is back but the damage is not fully repaired.

Current distribution

16 incidents

Closed638%
Monitoring531%
Actions open531%

Is policy and strategy coming to fruition

Whether the written intent is holding here

21,438 of 31,445 affected units have been reversed.

Not holding on the record

The written commitment is that no incident closes with unreversed work behind it. Against 16 incidents, 31,445 units of work were touched and 21,438 have been put right, leaving 10,007 outstanding. 5 incidents carry open remediation actions. Detection, containment and restore times on this page are as recorded by the responding team; none has been reviewed independently.