LensReading which lens this session carries.

Improve · item 33

Exception taxonomy and the learning loop

An exception handled well is a good day. The same exception handled well fourteen hundred times is a design failure nobody has named. This is the register of recurring failure classes: what each one is, how often it fires, who owns the fix, and what would have to be true to call it closed.

Policy sandbox

Health

Where automation hands work back

Exception classes by state and trend, with the case volume and handling time behind them.

Named classes
18
across 8 families · 0 closed
Occurrences in 90 days
9,389
up 2.8% on the prior ninety days
Handling hours
5,929
human hours consumed by these classes in the same window
No named owner
3
1,106 occurrences with nobody accountable for the fix
Root cause unknown
3
counted and contained, but not understood
Rising
10
7 falling · 2 reopened after being called fixed

By family

Grouping matters more than counting. Eight families, ranked by how much work they generate, with the share that has nobody accountable for the fix.

Intake2,385in 90 days2 classes · 809 hours
Model1,497in 90 days3 classes · 725 hours
Judgment1,393in 90 days2 classes · 1,021 hours
1 class in this family has no owner.
Matching1,055in 90 days2 classes · 847 hours
Human944in 90 days3 classes · 610 hours
1 class in this family has no owner.
Data853in 90 days2 classes · 914 hours
1 class in this family has no owner.
Policy778in 90 days2 classes · 463 hours
Handoff484in 90 days2 classes · 540 hours

The register

Ranked by ninety-day volume. Each class states its definition, whether the cause is understood, what the standard fix is, and the criterion that would let anyone call it closed.

3 with no standard fix · 2 with no closure criterion
Attachment the agent cannot readIntakeOpen1,418Rising · +18% on prior 90 days

A request arrives with an attachment the extraction step cannot parse: a photographed document, a password-protected file, or a spreadsheet whose figures live in a pivot cache rather than in cells.

Root cause

There is no upload contract. Any file type is accepted at intake and the failure surfaces three steps later, at which point the requester has moved on.

Standard fix

Validate the file at the point of upload and refuse it there with a specific reason, rather than accepting it and failing silently downstream.

Not deployed
Closed when

Fewer than 100 occurrences in a rolling 90 days, and no occurrence where the requester was not told at upload time.

Tomas Berger · Principal Engineer, Agent PlatformEstate-wide, concentrated in procurement and finance intake0.40 hours per case · 567 hours in 90 daysFirst seen 13 months ago · last seen today

This is the largest single exception class in the estate and it has been open for over a year. The fix is understood and unscheduled.

Model confidence below the acting thresholdModelOpen1,247Flat · +4% on prior 90 days

The agent produces an answer but its own confidence sits under the threshold at which it is permitted to act, so it hands the case to a person with its draft attached.

Root cause

The acting threshold is deliberately conservative and has not been revisited since the estate went live. It was set to be defensible rather than to be right.

Standard fix

Recalibrate the threshold per agent class against observed accuracy at each confidence band, rather than holding one figure across the whole estate.

Not deployed
Closed when

Per-class thresholds published with the accuracy evidence behind each one.

Marcus Weill · Director, Model GovernanceEstate-wide0.35 hours per case · 436 hours in 90 daysFirst seen 12 months ago · last seen today

Lowering this threshold would move a large volume off people and is exactly the kind of change that must go through the sandbox and an approval rather than being made quietly.

Request whose scope the agent cannot settleJudgmentOpen1,106Rising · +6% on prior 90 days

The request is legible but genuinely ambiguous, and the agent correctly refuses to guess. This is a designed refusal rather than a defect, and it is counted here so the volume of designed refusals is visible.

Root cause

The requester supplies less than the agent needs and the intake form does not ask for it, because asking for it lengthens the form.

Standard fix

A clarifying question returned within the same session rather than a routed exception, for the six request types that produce most of this volume.

Not deployed
Closed when

Half of these resolved by a clarifying question inside the session rather than by routing to a person.

Clara Vogt · Quality Lead, Agent EvaluationEstate-wide0.30 hours per case · 332 hours in 90 daysFirst seen 12 months ago · last seen today

Counting a designed refusal alongside genuine defects flatters the estate if it is read carelessly. It is separated by family here for exactly that reason.

Request submitted through the wrong channelIntakeContained967Falling · −14% on prior 90 days

Work arrives by email or chat to a named person rather than through the intake path the agents watch, so no agent ever sees it and the clock starts late.

Root cause

Long-standing habit reinforced by the fact that emailing a known person still works and is faster for the requester.

Standard fix

Auto-forward from the shared function mailboxes into intake, and a monthly note to the top twenty senders. Nothing is blocked, because blocking the channel pushes the work further underground.

Deployed May 14, 2026
Closed when

Fewer than 400 occurrences in a rolling 90 days for two consecutive quarters.

Avery Chen · Chief of StaffEstate-wide0.25 hours per case · 242 hours in 90 daysFirst seen 16 months ago · last seen yesterday

The count is falling but only from the mailboxes we know about. Direct messages to individuals are not measured at all, so the true number is higher than the one shown.

Invoice inside tolerance but outside patternMatchingOpen842Rising · +5% on prior 90 days

A three-way match falls within the price tolerance but the supplier, the cost center or the line description does not fit the historical pattern, so the agent refuses to auto-post and routes for review.

Root cause

The tolerance rule and the pattern check disagree by design. The tolerance was set by policy, the pattern check was learned from history, and nobody has reconciled the two.

Standard fix

Reconcile the pattern check against the published tolerance so the two do not contradict each other, and publish whichever is the binding rule.

Not deployed
Closed when

A single published matching rule, and fewer than 200 review routings in a rolling 90 days.

Marta Alvarez · Cost Accounting ManagerProcurement and finance shared services0.55 hours per case · 463 hours in 90 daysFirst seen 10 months ago · last seen todayReopened once after being called fixed

Closed once on the strength of a tolerance change, then reopened when the count went back up within six weeks. The first closure was called too early.

Master data disagrees between two systemsDataOpen731Rising · +6% on prior 90 days

The same entity carries different attributes in two systems the agent reads, and the agent stops rather than choosing which one is right.

Root cause

No system of record has been declared for the contested attributes. Each system believes it is authoritative.

Standard fix

Declare a system of record per attribute, publish it, and make the agents read only from it.

Not deployed
Closed when

A published system-of-record register covering every attribute the agents read.

No ownerEstate-wide0.90 hours per case · 658 hours in 90 daysFirst seen 21 months ago · last seen today

Unowned, second largest class by volume, and the fix is a governance decision rather than an engineering task. It has been described in three prior programs and never settled.

Policy threshold overtaken by the businessPolicyOpen604Rising · +18% on prior 90 days

An approval threshold set years ago now catches routine work, so the agent escalates volume that the policy never intended to catch.

Root cause

Thresholds are set in absolute currency and were never indexed. Prices moved and the thresholds did not.

Standard fix

Review the twelve highest-volume thresholds against current unit prices, and set a review date on each one rather than leaving them open-ended.

Not deployed
Closed when

Every threshold in the top twelve carries a review date, and escalation volume from stale thresholds falls below 300 in a rolling 90 days.

Helena Marsh · VP FP&AEstate-wide, worst in marketing and sales0.45 hours per case · 272 hours in 90 daysFirst seen 18 months ago · last seen today

This is the class the policy sandbox was built for. Nothing has yet been committed through it against this class.

Approval request expired without a decisionHumanOpen512Rising · +9% on prior 90 days

An agent routed a case for approval and no decision was made inside the window, so the case aged out and had to be resubmitted.

Root cause

Approvals are concentrated on a small number of people, and the platform routes by role without regard to how much is already in that person queue.

Standard fix

Route to the least loaded qualified approver rather than the first one, and escalate on the timeout rather than expiring the case.

Not deployed
Closed when

No case expires for want of a decision. Timeouts escalate instead of expiring.

Avery Chen · Chief of StaffEstate-wide0.80 hours per case · 410 hours in 90 daysFirst seen 10 months ago · last seen today

The current behavior loses work silently. Expiring a case rather than escalating it is a design fault, not a capacity problem.

Context lost between two agentsHandoffOpen388Rising · +9% on prior 90 days

Work passes from one agent to another and the second agent asks the requester for something the first already collected.

Root cause

Each agent holds its own working context and the handoff passes a case identifier rather than the collected evidence.

Standard fix

Pass the collected evidence with the handoff, and make re-asking a hard failure in the evaluation suite rather than a soft quality note.

Not deployed
Closed when

The re-ask case in the handoff evaluation suite passes for every agent in the chain, and the count falls below 100.

Tomas Berger · Principal Engineer, Agent PlatformCross-function chains, worst on hire-to-retire0.60 hours per case · 233 hours in 90 daysFirst seen 10 months ago · last seen today

The requester experiences this as the estate being disorganized, and it is the single most common complaint in the satisfaction comments.

Human override with no reason recordedHumanOpen344Falling · −8% on prior 90 days

A person overturned an agent decision and left the reason field empty, so the override cannot be fed back into evaluation and the agent cannot learn from it.

Root cause

The reason field is optional, and making it mandatory was rejected once on the grounds that it would slow approvers down.

Standard fix

A short structured reason with four options and a free-text box, mandatory on override only, not on approval.

Not deployed
Closed when

Over 90 percent of overrides carry a reason, and the reasons are read into the evaluation backlog monthly.

Clara Vogt · Quality Lead, Agent EvaluationEstate-wide0.20 hours per case · 69 hours in 90 daysFirst seen 11 months ago · last seen today

Every unexplained override is a lost training signal. This class is the reason the learning loop is thinner than it looks.

Two prior decisions point opposite waysJudgmentOpen287Rising · +18% on prior 90 days

The agent finds two comparable prior cases that were decided differently and has no basis to prefer one, so it escalates with both attached.

Root cause

Not established. The class is counted and contained, and nobody has found what causes it.

Standard fix

None written. Every occurrence is solved individually, which is why the count is not falling.

Not deployed
Closed when

Not defined. No owner has been named to define one.

No ownerLegal and commercial contracting2.40 hours per case · 689 hours in 90 daysFirst seen 9 months ago · last seen 2 days ago

Rising, unowned, and the most expensive class per case in the register. Legal is outside the lifecycle pilot, which is part of why nobody has picked it up.

Same supplier under two master recordsMatchingContained213Falling · −21% on prior 90 days

A supplier exists twice in the vendor master with different identifiers, so spend, terms and risk ratings split across two records and neither is complete.

Root cause

Two legacy vendor masters were merged in a prior program without a deduplication pass, and new duplicates keep being created because the create screen does not check for near matches.

Standard fix

A near-match check at create time, plus a monthly merge queue worked by the vendor master team.

Deployed Jun 25, 2026
Closed when

Fewer than 50 new duplicate pairs in a rolling 90 days and the historical backlog worked to zero.

Rafael Ortiz · Director, Source-to-PayVendor master, all regions1.80 hours per case · 383 hours in 90 daysFirst seen 23 months ago · last seen 3 days ago
Agent refused work it was permitted to doModelContained209Falling · −21% on prior 90 days

The agent declines a case that sits inside its published mandate. A false refusal is cheaper than a false action but it is still a defect, and it erodes trust faster than a slow answer.

Root cause

Refusal prompts were tightened after an early incident and were tightened further than the mandate requires.

Standard fix

A refusal case set in every evaluation suite that fails the agent for refusing inside its mandate, not only for acting outside it.

Deployed Jun 17, 2026
Closed when

Under 100 in a rolling 90 days with the refusal case set passing across the pilot.

Clara Vogt · Quality Lead, Agent EvaluationEstate-wide0.50 hours per case · 105 hours in 90 daysFirst seen 8 months ago · last seen 4 days ago

The refusal case set exists only for the pilot suites. Outside the pilot there is no test that would catch this.

Case in a jurisdiction with no mapped rulePolicyOpen174Falling · −7% on prior 90 days

Work arrives from an entity or country for which no policy variant has been loaded, so the agent applies the global default and flags that it did so.

Root cause

Policy variants were loaded for the markets with the most volume. The long tail was deferred and never revisited.

Standard fix

Load the variants for the next fifteen entities by volume, and make the global default refuse rather than proceed where the local rule is unknown.

Not deployed
Closed when

No case proceeds on a global default where a local rule exists but is unloaded.

Grace Abbott · Head of Internal AuditEntities outside the top nine markets1.10 hours per case · 191 hours in 90 daysFirst seen 14 months ago · last seen 5 days ago

The agent currently proceeds and flags. Audit has asked for it to refuse and wait, and that change has not been made.

Upstream feed arrives after the agent runsDataContained122Falling · −31% on prior 90 days

A scheduled agent runs against data that has not yet landed, produces an incomplete result, and the correction has to be made by hand.

Root cause

Agents run on a clock rather than on a data-ready signal, because two of the upstream feeds do not emit one.

Standard fix

Run on a data-ready signal where the feed emits one, and hold rather than proceed where it does not.

Deployed Jun 8, 2026
Closed when

No agent produces a result against an incomplete feed for two consecutive closes.

Ana Duarte · Group Financial ControllerRecord-to-report and month-end close2.10 hours per case · 256 hours in 90 daysFirst seen 13 months ago · last seen 9 days ago

Two feeds still have no ready signal, so the agents that depend on them still run on a clock.

Escalation with no one on the other sideHandoffFixed, not yet closed96Falling · −32% on prior 90 days

An agent escalates to a named role, and the person holding that role has left, changed team or is on extended leave with no delegate set.

Root cause

Escalation targets are held as names in the agent configuration rather than as roles resolved against the directory at the moment of escalation.

Standard fix

Resolve escalation targets against the directory at escalation time, and fall back to the role owner one level up when the seat is vacant.

Deployed Jul 11, 2026
Closed when

Zero escalations landing on a vacant seat over a full quarter.

Priya Raman · Head of Agent OperationsEstate-wide3.20 hours per case · 307 hours in 90 daysFirst seen 11 months ago · last seen 6 days agoReopened 2 times after being called fixed

Reopened twice. The fix has held for 38 days, which is shorter than the interval at which it previously failed, so it is shown as fixed rather than closed.

Work completed outside the platform entirelyHumanOpen88Rising · +44% on prior 90 days

A team completed work by hand that the agents were meant to handle, and the platform only learns about it when the record appears already finished.

Root cause

Not established. The class is counted and contained, and nobody has found what causes it.

Standard fix

None written. Every occurrence is solved individually, which is why the count is not falling.

Not deployed
Closed when

Not defined.

No ownerEstate-wide, identified by reconciliation rather than by report1.50 hours per case · 132 hours in 90 daysFirst seen 5 months ago · last seen 7 days ago

Rising, unowned, and detected only where a reconciliation happens to catch it. The count is a floor, not a measurement, and it is the class most likely to be understated on this page.

Output distribution drifted from the evaluated baselineModelOpen41Rising · +86% on prior 90 days

The monitoring layer flags that an agent output distribution has moved away from the one it was evaluated on, without any single case having failed.

Root cause

Not established. The class is counted and contained, and nobody has found what causes it.

Standard fix

None written. Every occurrence is solved individually, which is why the count is not falling.

Not deployed
Closed when

A drift flag is either explained by a known upstream change or it triggers a re-evaluation, with no flag sitting unexplained for more than fourteen days.

Marcus Weill · Director, Model GovernanceAgents inside the lifecycle pilot only4.50 hours per case · 185 hours in 90 daysFirst seen 4 months ago · last seen yesterday

Nearly doubled quarter on quarter, and the cause is unknown. Drift is only monitored for the pilot, so the estate-wide figure is not knowable from this platform.

Actions

What is waiting on a person

Exceptions are where automation quietly hands work back to people. These are the classes that cost the most.

Operations

What this desk is allowed to start

A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.

Trigger and bound

This desk can classify an exception, count its recurrence and price the handling time it consumes. It cannot change the upstream process, alter a rule, or suppress a class — fixing an exception class always means changing something on another page.

Live observability

What the record shows right now

Class state. Contained means the damage is limited but the cause is untouched.

Current distribution

18 classes

Fixed16%
Contained422%
Open1372%

Is policy and strategy coming to fruition

Whether the written intent is holding here

1 of 18 exception classes are actually fixed. 9,389 cases in 90 days cost about 5,929 hours of human handling.

Not holding on the record

The strategy position is that automation reduces manual handling. Exceptions are where that claim is tested, and it is not holding: only 1 of 18 classes are fixed, 10 are growing and 3 have no known cause. Handling time is calculated from the hours-per-case figure authored on each class multiplied by its 90-day count — it is a modeled cost, not a timesheet, and should be read as an order of magnitude.

What this page is, and what it is not

Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.

Naming a failure class is the cheap half. 18 classes are named and counted here, and 5 have had a standard fix deployed — 28%. 3 have no owner at all, covering 1,106 occurrences in the last ninety days, and 2 have no written criterion for what closed would even look like. A taxonomy without owners is a list, not a loop.

The oldest class here has been open for 1.9 years. Occurrences across the register are up 2.8% against the prior ninety days, which is the only honest headline: the loop is recording, and it is not yet reducing.