Improve · item 33
Exception taxonomy and the learning loop
An exception handled well is a good day. The same exception handled well fourteen hundred times is a design failure nobody has named. This is the register of recurring failure classes: what each one is, how often it fires, who owns the fix, and what would have to be true to call it closed.
Health
Where automation hands work back
Exception classes by state and trend, with the case volume and handling time behind them.
By family
Grouping matters more than counting. Eight families, ranked by how much work they generate, with the share that has nobody accountable for the fix.
The register
Ranked by ninety-day volume. Each class states its definition, whether the cause is understood, what the standard fix is, and the criterion that would let anyone call it closed.
A request arrives with an attachment the extraction step cannot parse: a photographed document, a password-protected file, or a spreadsheet whose figures live in a pivot cache rather than in cells.
There is no upload contract. Any file type is accepted at intake and the failure surfaces three steps later, at which point the requester has moved on.
Validate the file at the point of upload and refuse it there with a specific reason, rather than accepting it and failing silently downstream.
Fewer than 100 occurrences in a rolling 90 days, and no occurrence where the requester was not told at upload time.
This is the largest single exception class in the estate and it has been open for over a year. The fix is understood and unscheduled.
The agent produces an answer but its own confidence sits under the threshold at which it is permitted to act, so it hands the case to a person with its draft attached.
The acting threshold is deliberately conservative and has not been revisited since the estate went live. It was set to be defensible rather than to be right.
Recalibrate the threshold per agent class against observed accuracy at each confidence band, rather than holding one figure across the whole estate.
Per-class thresholds published with the accuracy evidence behind each one.
Lowering this threshold would move a large volume off people and is exactly the kind of change that must go through the sandbox and an approval rather than being made quietly.
The request is legible but genuinely ambiguous, and the agent correctly refuses to guess. This is a designed refusal rather than a defect, and it is counted here so the volume of designed refusals is visible.
The requester supplies less than the agent needs and the intake form does not ask for it, because asking for it lengthens the form.
A clarifying question returned within the same session rather than a routed exception, for the six request types that produce most of this volume.
Half of these resolved by a clarifying question inside the session rather than by routing to a person.
Counting a designed refusal alongside genuine defects flatters the estate if it is read carelessly. It is separated by family here for exactly that reason.
Work arrives by email or chat to a named person rather than through the intake path the agents watch, so no agent ever sees it and the clock starts late.
Long-standing habit reinforced by the fact that emailing a known person still works and is faster for the requester.
Auto-forward from the shared function mailboxes into intake, and a monthly note to the top twenty senders. Nothing is blocked, because blocking the channel pushes the work further underground.
Fewer than 400 occurrences in a rolling 90 days for two consecutive quarters.
The count is falling but only from the mailboxes we know about. Direct messages to individuals are not measured at all, so the true number is higher than the one shown.
A three-way match falls within the price tolerance but the supplier, the cost center or the line description does not fit the historical pattern, so the agent refuses to auto-post and routes for review.
The tolerance rule and the pattern check disagree by design. The tolerance was set by policy, the pattern check was learned from history, and nobody has reconciled the two.
Reconcile the pattern check against the published tolerance so the two do not contradict each other, and publish whichever is the binding rule.
A single published matching rule, and fewer than 200 review routings in a rolling 90 days.
Closed once on the strength of a tolerance change, then reopened when the count went back up within six weeks. The first closure was called too early.
The same entity carries different attributes in two systems the agent reads, and the agent stops rather than choosing which one is right.
No system of record has been declared for the contested attributes. Each system believes it is authoritative.
Declare a system of record per attribute, publish it, and make the agents read only from it.
A published system-of-record register covering every attribute the agents read.
Unowned, second largest class by volume, and the fix is a governance decision rather than an engineering task. It has been described in three prior programs and never settled.
An approval threshold set years ago now catches routine work, so the agent escalates volume that the policy never intended to catch.
Thresholds are set in absolute currency and were never indexed. Prices moved and the thresholds did not.
Review the twelve highest-volume thresholds against current unit prices, and set a review date on each one rather than leaving them open-ended.
Every threshold in the top twelve carries a review date, and escalation volume from stale thresholds falls below 300 in a rolling 90 days.
This is the class the policy sandbox was built for. Nothing has yet been committed through it against this class.
An agent routed a case for approval and no decision was made inside the window, so the case aged out and had to be resubmitted.
Approvals are concentrated on a small number of people, and the platform routes by role without regard to how much is already in that person queue.
Route to the least loaded qualified approver rather than the first one, and escalate on the timeout rather than expiring the case.
No case expires for want of a decision. Timeouts escalate instead of expiring.
The current behavior loses work silently. Expiring a case rather than escalating it is a design fault, not a capacity problem.
Work passes from one agent to another and the second agent asks the requester for something the first already collected.
Each agent holds its own working context and the handoff passes a case identifier rather than the collected evidence.
Pass the collected evidence with the handoff, and make re-asking a hard failure in the evaluation suite rather than a soft quality note.
The re-ask case in the handoff evaluation suite passes for every agent in the chain, and the count falls below 100.
The requester experiences this as the estate being disorganized, and it is the single most common complaint in the satisfaction comments.
A person overturned an agent decision and left the reason field empty, so the override cannot be fed back into evaluation and the agent cannot learn from it.
The reason field is optional, and making it mandatory was rejected once on the grounds that it would slow approvers down.
A short structured reason with four options and a free-text box, mandatory on override only, not on approval.
Over 90 percent of overrides carry a reason, and the reasons are read into the evaluation backlog monthly.
Every unexplained override is a lost training signal. This class is the reason the learning loop is thinner than it looks.
The agent finds two comparable prior cases that were decided differently and has no basis to prefer one, so it escalates with both attached.
Not established. The class is counted and contained, and nobody has found what causes it.
None written. Every occurrence is solved individually, which is why the count is not falling.
Not defined. No owner has been named to define one.
Rising, unowned, and the most expensive class per case in the register. Legal is outside the lifecycle pilot, which is part of why nobody has picked it up.
A supplier exists twice in the vendor master with different identifiers, so spend, terms and risk ratings split across two records and neither is complete.
Two legacy vendor masters were merged in a prior program without a deduplication pass, and new duplicates keep being created because the create screen does not check for near matches.
A near-match check at create time, plus a monthly merge queue worked by the vendor master team.
Fewer than 50 new duplicate pairs in a rolling 90 days and the historical backlog worked to zero.
The agent declines a case that sits inside its published mandate. A false refusal is cheaper than a false action but it is still a defect, and it erodes trust faster than a slow answer.
Refusal prompts were tightened after an early incident and were tightened further than the mandate requires.
A refusal case set in every evaluation suite that fails the agent for refusing inside its mandate, not only for acting outside it.
Under 100 in a rolling 90 days with the refusal case set passing across the pilot.
The refusal case set exists only for the pilot suites. Outside the pilot there is no test that would catch this.
Work arrives from an entity or country for which no policy variant has been loaded, so the agent applies the global default and flags that it did so.
Policy variants were loaded for the markets with the most volume. The long tail was deferred and never revisited.
Load the variants for the next fifteen entities by volume, and make the global default refuse rather than proceed where the local rule is unknown.
No case proceeds on a global default where a local rule exists but is unloaded.
The agent currently proceeds and flags. Audit has asked for it to refuse and wait, and that change has not been made.
A scheduled agent runs against data that has not yet landed, produces an incomplete result, and the correction has to be made by hand.
Agents run on a clock rather than on a data-ready signal, because two of the upstream feeds do not emit one.
Run on a data-ready signal where the feed emits one, and hold rather than proceed where it does not.
No agent produces a result against an incomplete feed for two consecutive closes.
Two feeds still have no ready signal, so the agents that depend on them still run on a clock.
An agent escalates to a named role, and the person holding that role has left, changed team or is on extended leave with no delegate set.
Escalation targets are held as names in the agent configuration rather than as roles resolved against the directory at the moment of escalation.
Resolve escalation targets against the directory at escalation time, and fall back to the role owner one level up when the seat is vacant.
Zero escalations landing on a vacant seat over a full quarter.
Reopened twice. The fix has held for 38 days, which is shorter than the interval at which it previously failed, so it is shown as fixed rather than closed.
A team completed work by hand that the agents were meant to handle, and the platform only learns about it when the record appears already finished.
Not established. The class is counted and contained, and nobody has found what causes it.
None written. Every occurrence is solved individually, which is why the count is not falling.
Not defined.
Rising, unowned, and detected only where a reconciliation happens to catch it. The count is a floor, not a measurement, and it is the class most likely to be understated on this page.
The monitoring layer flags that an agent output distribution has moved away from the one it was evaluated on, without any single case having failed.
Not established. The class is counted and contained, and nobody has found what causes it.
None written. Every occurrence is solved individually, which is why the count is not falling.
A drift flag is either explained by a known upstream change or it triggers a re-evaluation, with no flag sitting unexplained for more than fourteen days.
Nearly doubled quarter on quarter, and the cause is unknown. Drift is only monitored for the pilot, so the estate-wide figure is not knowable from this platform.
Actions
What is waiting on a person
Exceptions are where automation quietly hands work back to people. These are the classes that cost the most.
Own an open exception class
No containment and no fix in progress.
Attack a class that is growing
More cases this quarter than last. It gets worse if left.
Convert a containment into a fix
The symptom is managed. The cause is still there.
Find the root cause of a class
Handled repeatedly with no cause identified.
Operations
What this desk is allowed to start
A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.
Trigger and bound
This desk can classify an exception, count its recurrence and price the handling time it consumes. It cannot change the upstream process, alter a rule, or suppress a class — fixing an exception class always means changing something on another page.
Live observability
What the record shows right now
Class state. Contained means the damage is limited but the cause is untouched.
Current distribution
18 classes
Is policy and strategy coming to fruition
Whether the written intent is holding here
1 of 18 exception classes are actually fixed. 9,389 cases in 90 days cost about 5,929 hours of human handling.
Not holding on the record
The strategy position is that automation reduces manual handling. Exceptions are where that claim is tested, and it is not holding: only 1 of 18 classes are fixed, 10 are growing and 3 have no known cause. Handling time is calculated from the hours-per-case figure authored on each class multiplied by its 90-day count — it is a modeled cost, not a timesheet, and should be read as an order of magnitude.
Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.
Naming a failure class is the cheap half. 18 classes are named and counted here, and 5 have had a standard fix deployed — 28%. 3 have no owner at all, covering 1,106 occurrences in the last ninety days, and 2 have no written criterion for what closed would even look like. A taxonomy without owners is a list, not a loop.
The oldest class here has been open for 1.9 years. Occurrences across the register are up 2.8% against the prior ninety days, which is the only honest headline: the loop is recording, and it is not yet reducing.