LensReading which lens this session carries.

Improve · item 29

Evaluation harness

A promotion is only auditable if the evidence behind it survives the promotion. Every suite here is a fixed set of cases with a written expected answer, scored by version, so a reviewer can come back a year later and see exactly what the agent could do on the day it was trusted with more.

Lifecycle

Health

Whether the agents are still being tested

Suites by health, and the pass, fail and regression verdicts of the runs behind them.

Suites
115
covering 115 of 178 managed agents
Cases written
1,308
11.4 per suite on average · 26 under nine cases
Mean latest score
90.8%
103 cases failing · 44 skipped
Runs attached to a promotion
191
of 460 runs — the rest are routine regression passes
Independently observed
97 / 460
the rest were run and scored by the team that built the agent
Agents with no suite
63
nothing to test them against, so nothing can promote them
Read the harness honestly
30
of 115 suites hold nothing back. Every case is visible to the people tuning the agent, so a high score can be reached by fitting the test.
37
suites never measured whether two humans agree on the expected answer. Where labelers disagree, the score measures the labeler, not the agent.
71
suites contain no adversarial cases at all. They test the agent on work that arrives in good faith and say nothing about work that does not.
92
suites hold no matched pairs for disparate impact. A suite without them cannot detect a decision that is consistently right and consistently unfair.

Every suite, and what it does not cover

Coverage is stated as a limit rather than a percentage, because a percentage of a set we wrote ourselves is not a measure of anything.

115 suites
Service Desk Orchestrator — golden setThin coverage100.0%6 cases · 4 runs · last 8 days ago

Cases the service desk orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner5 human labeled89% agreementHeld-out split kept backOwner Tomas Berger

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Intake & Classification Agent — golden setOpen failures85.7%16 cases · 4 runs · last 19 days ago

Cases the intake & classification agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Ibrahim Sy

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Tier-1 Resolution Agent — golden setThin coverage100.0%8 cases · 4 runs · last 2 months ago

Cases the tier-1 resolution agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim7 human labeledAgreement not measuredNo held-out splitOwner Ibrahim Sy

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Knowledge Retrieval Agent — golden setThin coverage100.0%8 cases · 4 runs · last 2 months ago

Cases the knowledge retrieval agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim7 human labeledAgreement not measuredHeld-out split kept backOwner Elena Farrow

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Escalation Routing Agent — golden setOpen failures62.5%16 cases · 4 runs · last 2 months ago

Cases the escalation routing agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim13 human labeled91% agreementHeld-out split kept backOwner Clara Vogt

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollback
Service Desk Challenger Agent — golden setOpen failures71.4%14 cases · 4 runs · last 17 days ago

Cases the service desk challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification11 human labeled90% agreementHeld-out split kept backOwner Nadia Kovac

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Data Stewardship Orchestrator — golden setPassing100.0%12 cases · 4 runs · last 41 days ago

Cases the data stewardship orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated11 human labeled97% agreementHeld-out split kept backOwner Nadia Kovac

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Duplicate Detection Agent — golden setOpen failures75.0%16 cases · 4 runs · last 26 days ago

Cases the duplicate detection agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong14 human labeled95% agreementHeld-out split kept backOwner Tomas Berger

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Enrichment Agent — golden setPassing90.0%10 cases · 4 runs · last 2 months ago

Cases the enrichment agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated9 human labeled95% agreementHeld-out split kept backOwner Clara Vogt

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Ownership Reconciliation Agent — golden setPassing100.0%12 cases · 4 runs · last 14 days ago

Cases the ownership reconciliation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated10 human labeledAgreement not measuredHeld-out split kept backOwner Marcus Weill

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Data Policy Agent — golden setOpen failures83.3%12 cases · 4 runs · last 2 months ago

Cases the data policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim9 human labeledAgreement not measuredNo held-out splitOwner Tomas Berger

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at all
CRM Data Stewardship Challenger Agent — golden setThin coverage100.0%8 cases · 4 runs · last 14 days ago

Cases the crm data stewardship challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification6 human labeled90% agreementHeld-out split kept backOwner Ibrahim Sy

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Identity Orchestrator — golden setOpen failures73.3%16 cases · 4 runs · last 2 months ago

Cases the identity orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification12 human labeled87% agreementNo held-out splitOwner Daniel Okoye

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Joiner Provisioning Agent — golden setOpen failures78.6%14 cases · 4 runs · last 3 months ago

Cases the joiner provisioning agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim10 human labeledAgreement not measuredHeld-out split kept backOwner Clara Vogt

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Entitlement Review Agent — golden setPassing100.0%12 cases · 4 runs · last 2 months ago

Cases the entitlement review agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled80% agreementHeld-out split kept backOwner Nadia Kovac

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Privileged Access Agent — golden setPassing100.0%14 cases · 4 runs · last 2 months ago

Cases the privileged access agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner10 human labeled86% agreementHeld-out split kept backOwner Elena Farrow

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Duty Conflict Agent — golden setNo held-out set100.0%12 cases · 4 runs · last 3 months ago

Cases the duty conflict agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Access Policy Agent — golden setPassing100.0%12 cases · 4 runs · last 28 days ago

Cases the access policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong9 human labeled86% agreementHeld-out split kept backOwner Tomas Berger

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Price List Maintenance Agent — golden setThin coverage100.0%8 cases · 4 runs · last 23 days ago

Cases the price list maintenance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification5 human labeled84% agreementHeld-out split kept backOwner Nadia Kovac

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Competitive Price Watch Agent — golden setOpen failures68.8%16 cases · 4 runs · last 3 months ago

Cases the competitive price watch agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong10 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Discount Guardrail Agent — golden setThin coverage100.0%8 cases · 4 runs · last 8 days ago

Cases the discount guardrail agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification6 human labeled91% agreementHeld-out split kept backOwner Daniel Okoye

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Pricing Strategy Agent — golden setPassing100.0%10 cases · 4 runs · last 38 days ago

Cases the pricing strategy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled83% agreementHeld-out split kept backOwner Ibrahim Sy

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
End-User Compute Orchestrator — golden setOpen failures90.0%10 cases · 4 runs · last 41 days ago

Cases the end-user compute orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated6 human labeledAgreement not measuredNo held-out splitOwner Tomas Berger

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo disparate-impact pairs in this suite at all
Device Provisioning Agent — golden setThin coverage100.0%8 cases · 4 runs · last 6 days ago

Cases the device provisioning agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated5 human labeled80% agreementHeld-out split kept backOwner Daniel Okoye

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Patch Compliance Agent — golden setPassing83.3%12 cases · 4 runs · last 2 months ago

Cases the patch compliance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated9 human labeled87% agreementHeld-out split kept backOwner Marcus Weill

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Device Health Agent — golden setPassing100.0%10 cases · 4 runs · last 23 days ago

Cases the device health agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim9 human labeledAgreement not measuredHeld-out split kept backOwner Elena Farrow

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Refresh Planning Agent — golden setPassing100.0%10 cases · 4 runs · last 2 months ago

Cases the refresh planning agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeled87% agreementHeld-out split kept backOwner Nadia Kovac

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Device Policy Agent — golden setThin coverage100.0%6 cases · 4 runs · last 3 months ago

Cases the device policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner5 human labeled90% agreementHeld-out split kept backOwner Priya Raman

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
End-User Compute Challenger Agent — golden setNo held-out set66.7%12 cases · 4 runs · last 2 months ago

Cases the end-user compute challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong8 human labeled82% agreementNo held-out splitOwner Elena Farrow

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Quote Assembly Agent — golden setOpen failures90.0%10 cases · 4 runs · last 21 days ago

Cases the quote assembly agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeled78% agreementHeld-out split kept backOwner Elena Farrow

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Approval Routing Agent — golden setPassing100.0%12 cases · 4 runs · last 36 days ago

Cases the approval routing agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeled96% agreementHeld-out split kept backOwner Clara Vogt

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo disparate-impact pairs in this suite at all
Signature Coordination Agent — golden setOpen failures91.7%12 cases · 4 runs · last 18 days ago

Cases the signature coordination agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeled85% agreementHeld-out split kept backOwner Clara Vogt

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo disparate-impact pairs in this suite at all
Quote-to-Order Challenger Agent — golden setThin coverage50.0%8 cases · 4 runs · last 2 months ago

Cases the quote-to-order challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeled95% agreementHeld-out split kept backOwner Ibrahim Sy

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Order Validation Agent — golden setNo held-out set100.0%12 cases · 4 runs · last 8 days ago

Cases the order validation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeled81% agreementNo held-out splitOwner Daniel Okoye

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Provisioning Trigger Agent — golden setThin coverage100.0%8 cases · 4 runs · last 34 days ago

Cases the provisioning trigger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner6 human labeled86% agreementHeld-out split kept backOwner Daniel Okoye

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Order Exception Agent — golden setOpen failures100.0%14 cases · 4 runs · last 3 months ago

Cases the order exception agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim9 human labeled81% agreementNo held-out splitOwner Daniel Okoye

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Booking & Order Management Challenger Agent — golden setOpen failures84.6%14 cases · 4 runs · last 18 days ago

Cases the booking & order management challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeled80% agreementHeld-out split kept backOwner Marcus Weill

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Infrastructure Orchestrator — golden setNo held-out set100.0%14 cases · 4 runs · last 36 days ago

Cases the infrastructure orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification12 human labeledAgreement not measuredNo held-out splitOwner Nadia Kovac

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo disparate-impact pairs in this suite at all
Capacity Agent — golden setPassing100.0%10 cases · 4 runs · last 2 months ago

Cases the capacity agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeled97% agreementHeld-out split kept backOwner Ibrahim Sy

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Infrastructure Policy Agent — golden setNo held-out set100.0%14 cases · 4 runs · last 3 months ago

Cases the infrastructure policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated9 human labeled81% agreementNo held-out splitOwner Ibrahim Sy

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Infrastructure & Cloud Challenger Agent — golden setNo held-out set100.0%10 cases · 4 runs · last 44 days ago

Cases the infrastructure & cloud challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeled92% agreementNo held-out splitOwner Marcus Weill

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Billing Interlock Orchestrator — golden setOpen failures90.0%10 cases · 4 runs · last 31 days ago

Cases the billing interlock orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeled90% agreementHeld-out split kept backOwner Marcus Weill

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Invoice Accuracy Agent — golden setOpen failures75.0%8 cases · 4 runs · last 2 months ago

Cases the invoice accuracy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeledAgreement not measuredHeld-out split kept backOwner Nadia Kovac

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Billing Exception Agent — golden setOpen failures66.7%12 cases · 4 runs · last 16 days ago

Cases the billing exception agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification9 human labeled90% agreementHeld-out split kept backOwner Elena Farrow

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Billing Interlock Challenger Agent — golden setPassing100.0%10 cases · 4 runs · last 3 months ago

Cases the billing interlock challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Network Orchestrator — golden setPassing100.0%10 cases · 4 runs · last 3 months ago

Cases the network orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim9 human labeled94% agreementHeld-out split kept backOwner Tomas Berger

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Performance Agent — golden setPassing100.0%12 cases · 4 runs · last 15 days ago

Cases the performance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong8 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Configuration Agent — golden setOpen failures83.3%12 cases · 4 runs · last 2 months ago

Cases the configuration agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim8 human labeled84% agreementHeld-out split kept backOwner Nadia Kovac

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Segmentation Agent — golden setOpen failures60.0%10 cases · 4 runs · last 2 months ago

Cases the segmentation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeled78% agreementHeld-out split kept backOwner Ibrahim Sy

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Network Policy Agent — golden setPassing75.0%12 cases · 4 runs · last 3 months ago

Cases the network policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled81% agreementHeld-out split kept backOwner Marcus Weill

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Network Challenger Agent — golden setPassing100.0%12 cases · 4 runs · last 43 days ago

Cases the network challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner11 human labeled96% agreementHeld-out split kept backOwner Elena Farrow

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Application Management Orchestrator — golden setPassing100.0%10 cases · 4 runs · last 2 months ago

Cases the application management orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim7 human labeled87% agreementHeld-out split kept backOwner Ibrahim Sy

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Application Health Agent — golden setPassing100.0%10 cases · 4 runs · last 2 months ago

Cases the application health agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim6 human labeled79% agreementHeld-out split kept backOwner Elena Farrow

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Integration Monitoring Agent — golden setPassing100.0%10 cases · 4 runs · last 11 days ago

Cases the integration monitoring agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Vendor Escalation Agent — golden setPassing100.0%10 cases · 4 runs · last 19 days ago

Cases the vendor escalation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeledAgreement not measuredHeld-out split kept backOwner Daniel Okoye

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Configuration Agent — golden setOpen failures87.5%8 cases · 4 runs · last 37 days ago

Cases the configuration agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim6 human labeledAgreement not measuredNo held-out splitOwner Tomas Berger

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Application Management Challenger Agent — golden setNo held-out set100.0%12 cases · 4 runs · last 2 months ago

Cases the application management challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at all
Revenue Recognition Orchestrator — golden setOpen failures91.7%12 cases · 4 runs · last 28 days ago

Cases the revenue recognition orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner11 human labeled97% agreementHeld-out split kept backOwner Nadia Kovac

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Performance Obligation Agent — golden setPassing100.0%14 cases · 4 runs · last 3 months ago

Cases the performance obligation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim11 human labeled90% agreementHeld-out split kept backOwner Priya Raman

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Allocation Agent — golden setOpen failures87.5%8 cases · 4 runs · last 41 days ago

Cases the allocation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification7 human labeled96% agreementHeld-out split kept backOwner Daniel Okoye

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Deferred Balance Agent — golden setPassing100.0%10 cases · 4 runs · last 34 days ago

Cases the deferred balance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim8 human labeled88% agreementHeld-out split kept backOwner Daniel Okoye

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Audit Evidence Agent — golden setOpen failures80.0%10 cases · 4 runs · last 3 months ago

Cases the audit evidence agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification7 human labeled87% agreementHeld-out split kept backOwner Nadia Kovac

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Renewal Ops Orchestrator — golden setPassing100.0%14 cases · 4 runs · last 35 days ago

Cases the renewal ops orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated11 human labeled90% agreementHeld-out split kept backOwner Daniel Okoye

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo disparate-impact pairs in this suite at all
Renewal Calendar Agent — golden setPassing100.0%12 cases · 4 runs · last 42 days ago

Cases the renewal calendar agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification9 human labeled88% agreementHeld-out split kept backOwner Ibrahim Sy

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cancellation Agent — golden setPassing70.0%12 cases · 4 runs · last 3 months ago

Cases the cancellation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification11 human labeled96% agreementHeld-out split kept backOwner Priya Raman

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Retention Analytics Agent — golden setThin coverage100.0%8 cases · 4 runs · last 20 days ago

Cases the retention analytics agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim5 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Security Operations Orchestrator — golden setOpen failures100.0%16 cases · 4 runs · last 3 months ago

Cases the security operations orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeled85% agreementHeld-out split kept backOwner Ibrahim Sy

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibrated
Alert Triage Agent — golden setNo held-out set100.0%10 cases · 4 runs · last 2 months ago

Cases the alert triage agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification8 human labeled91% agreementNo held-out splitOwner Marcus Weill

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Threat Intelligence Agent — golden setPassing100.0%10 cases · 4 runs · last 40 days ago

Cases the threat intelligence agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Forensics Support Agent — golden setThin coverage87.5%8 cases · 4 runs · last 3 months ago

Cases the forensics support agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeled92% agreementNo held-out splitOwner Priya Raman

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Phishing Response Agent — golden setOpen failures85.7%16 cases · 4 runs · last 36 days ago

Cases the phishing response agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong12 human labeledAgreement not measuredNo held-out splitOwner Elena Farrow

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Vulnerability Agent — golden setOpen failures87.5%16 cases · 4 runs · last 2 months ago

Cases the vulnerability agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibrated
Security Policy Agent — golden setNo held-out set100.0%14 cases · 4 runs · last 2 months ago

Cases the security policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Marcus Weill

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Territory Ops Orchestrator — golden setThin coverage100.0%6 cases · 4 runs · last 43 days ago

Cases the territory ops orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong3 human labeled78% agreementHeld-out split kept backOwner Nadia Kovac

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Crediting Rules Agent — golden setOpen failures60.0%16 cases · 4 runs · last 31 days ago

Cases the crediting rules agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification12 human labeledAgreement not measuredHeld-out split kept backOwner Daniel Okoye

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Split Resolution Agent — golden setNo held-out set100.0%10 cases · 4 runs · last 34 days ago

Cases the split resolution agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeledAgreement not measuredNo held-out splitOwner Clara Vogt

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Quota Tracking Agent — golden setThin coverage100.0%8 cases · 4 runs · last 2 months ago

Cases the quota tracking agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner6 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Comp Ops Policy Agent — golden setOpen failures71.4%16 cases · 4 runs · last 25 days ago

Cases the comp ops policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated14 human labeled93% agreementNo held-out splitOwner Tomas Berger

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Territory & Comp Ops Challenger Agent — golden setThin coverage100.0%8 cases · 4 runs · last 18 days ago

Cases the territory & comp ops challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim6 human labeled89% agreementHeld-out split kept backOwner Clara Vogt

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at all
Data Platform Orchestrator — golden setPassing100.0%10 cases · 4 runs · last 2 months ago

Cases the data platform orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim7 human labeled83% agreementHeld-out split kept backOwner Marcus Weill

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Pipeline Health Agent — golden setPassing100.0%14 cases · 4 runs · last 22 days ago

Cases the pipeline health agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeled90% agreementHeld-out split kept backOwner Tomas Berger

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo disparate-impact pairs in this suite at all
Data Quality Agent — golden setPassing100.0%12 cases · 4 runs · last 15 days ago

Cases the data quality agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner10 human labeled93% agreementHeld-out split kept backOwner Nadia Kovac

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Access Governance Agent — golden setOpen failures71.4%14 cases · 4 runs · last 39 days ago

Cases the access governance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Forecast Integrity Orchestrator — golden setNo held-out set85.7%14 cases · 4 runs · last 3 months ago

Cases the forecast integrity orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong9 human labeledAgreement not measuredNo held-out splitOwner Marcus Weill

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo disparate-impact pairs in this suite at all
Submission Compliance Agent — golden setOpen failures87.5%16 cases · 4 runs · last 2 months ago

Cases the submission compliance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Clara Vogt

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollback
Model vs Commit Agent — golden setPassing83.3%12 cases · 4 runs · last 3 months ago

Cases the model vs commit agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Optimism Bias Agent — golden setThin coverage100.0%8 cases · 4 runs · last 28 days ago

Cases the optimism bias agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner4 human labeledAgreement not measuredHeld-out split kept backOwner Marcus Weill

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Integrity Policy Agent — golden setPassing100.0%10 cases · 4 runs · last 3 months ago

Cases the integrity policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled83% agreementHeld-out split kept backOwner Tomas Berger

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Forecast Integrity Challenger Agent — golden setPassing100.0%14 cases · 4 runs · last 2 months ago

Cases the forecast integrity challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner11 human labeled90% agreementHeld-out split kept backOwner Priya Raman

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo disparate-impact pairs in this suite at all
Discovery Agent — golden setThin coverage62.5%8 cases · 4 runs · last 2 months ago

Cases the discovery agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong5 human labeled83% agreementHeld-out split kept backOwner Ibrahim Sy

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
License Reconciliation Agent — golden setThin coverage100.0%8 cases · 4 runs · last 39 days ago

Cases the license reconciliation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim7 human labeled94% agreementHeld-out split kept backOwner Tomas Berger

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Renewal Agent — golden setPassing100.0%12 cases · 4 runs · last 9 days ago

Cases the renewal agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim10 human labeled93% agreementHeld-out split kept backOwner Priya Raman

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Utilization Agent — golden setPassing100.0%10 cases · 4 runs · last 19 days ago

Cases the utilization agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeled79% agreementHeld-out split kept backOwner Elena Farrow

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Reharvest Agent — golden setPassing100.0%12 cases · 4 runs · last 42 days ago

Cases the reharvest agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeled79% agreementHeld-out split kept backOwner Nadia Kovac

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Asset Policy Agent — golden setPassing100.0%12 cases · 4 runs · last 38 days ago

Cases the asset policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeledAgreement not measuredHeld-out split kept backOwner Clara Vogt

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Revenue Systems Orchestrator — golden setNo held-out set100.0%12 cases · 4 runs · last 3 months ago

Cases the revenue systems orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeled86% agreementNo held-out splitOwner Ibrahim Sy

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Integration Health Monitor — golden setPassing100.0%10 cases · 4 runs · last 29 days ago

Cases the integration health monitor must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Sandbox Refresh Agent — golden setOpen failures58.3%12 cases · 4 runs · last 2 months ago

Cases the sandbox refresh agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong9 human labeled90% agreementHeld-out split kept backOwner Marcus Weill

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Systems Policy Agent — golden setOpen failures75.0%16 cases · 4 runs · last 2 months ago

Cases the systems policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong13 human labeled91% agreementHeld-out split kept backOwner Daniel Okoye

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Systems & Integration Challenger Agent — golden setPassing100.0%12 cases · 4 runs · last 21 days ago

Cases the systems & integration challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification8 human labeled82% agreementHeld-out split kept backOwner Clara Vogt

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Risk Scoring Agent — golden setThin coverage100.0%8 cases · 4 runs · last 2 months ago

Cases the risk scoring agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner6 human labeled90% agreementHeld-out split kept backOwner Ibrahim Sy

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Change Board Prep Agent — golden setOpen failures75.0%8 cases · 4 runs · last 2 months ago

Cases the change board prep agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim6 human labeled89% agreementHeld-out split kept backOwner Nadia Kovac

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Change & Release Challenger Agent — golden setOpen failures85.7%14 cases · 4 runs · last 12 days ago

Cases the change & release challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeled82% agreementNo held-out splitOwner Daniel Okoye

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
IT Finance Orchestrator — golden setPassing100.0%12 cases · 4 runs · last 28 days ago

Cases the it finance orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Production exceptions from the last four quarters, retained verbatim8 human labeled85% agreementHeld-out split kept backOwner Priya Raman

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Cost Attribution Agent — golden setNo held-out set100.0%10 cases · 4 runs · last 26 days ago

Cases the cost attribution agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification6 human labeled82% agreementNo held-out splitOwner Nadia Kovac

The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.

No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Showback Agent — golden setThin coverage62.5%8 cases · 4 runs · last 2 months ago

Cases the showback agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeledAgreement not measuredHeld-out split kept backOwner Clara Vogt

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Budget Tracking Agent — golden setThin coverage100.0%8 cases · 4 runs · last 21 days ago

Cases the budget tracking agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong5 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman

The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.

Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Unit Cost Agent — golden setOpen failures91.7%12 cases · 4 runs · last 2 months ago

Cases the unit cost agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeledAgreement not measuredNo held-out splitOwner Clara Vogt

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight points
Sourcing Model Agent — golden setOpen failures91.7%12 cases · 4 runs · last 2 months ago

Cases the sourcing model agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification10 human labeled94% agreementHeld-out split kept backOwner Clara Vogt

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo disparate-impact pairs in this suite at all
IT Financial Management Challenger Agent — golden setNo held-out set100.0%14 cases · 4 runs · last 8 days ago

Cases the it financial management challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification11 human labeledAgreement not measuredNo held-out splitOwner Ibrahim Sy

Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.

No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo disparate-impact pairs in this suite at all
Revenue Analytics Orchestrator — golden setOpen failures85.7%16 cases · 4 runs · last 18 days ago

Cases the revenue analytics orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated13 human labeled90% agreementHeld-out split kept backOwner Daniel Okoye

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Cohort Analysis Agent — golden setOpen failures75.0%8 cases · 4 runs · last 25 days ago

Cases the cohort analysis agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong8 human labeled97% agreementHeld-out split kept backOwner Daniel Okoye

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Retention Modeling Agent — golden setOpen failures87.5%16 cases · 4 runs · last 37 days ago

Cases the retention modeling agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Cases written by the process owner before the agent existed, as a specification13 human labeled92% agreementHeld-out split kept backOwner Daniel Okoye

Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.

Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Pipeline Velocity Agent — golden setNo held-out set83.3%14 cases · 4 runs · last 3 months ago

Cases the pipeline velocity agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated13 human labeled97% agreementNo held-out splitOwner Marcus Weill

Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.

No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo disparate-impact pairs in this suite at all
Revenue Architecture Agent — golden setOpen failures80.0%12 cases · 4 runs · last 40 days ago

Cases the revenue architecture agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.

Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeled84% agreementHeld-out split kept backOwner Clara Vogt

Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.

No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at all

What the cases actually test

Cases by category across every suite, with the count currently failing.

happy-path23024 failing
present in 115 suites
edge22420 failing
present in 112 suites
regression22418 failing
present in 112 suites
ambiguity18813 failing
present in 94 suites
policy16218 failing
present in 81 suites
refusal1465 failing
present in 73 suites
adversarial883 failing
present in 44 suites
fairness462 failing
present in 23 suites

Regressions caught by version

A run is a regression when a new agent version scores below the version before it, whatever the absolute score.

23 of 460
v1.4100.0% → 91.7%Jul 31, 2026
12 cases executed against version 1.4. 1 case failed.
Cases: C-12
v1.4100.0% → 90.0%Jul 28, 2026
10 cases executed against version 1.4. 1 case failed.
Cases: C-01
v1.4100.0% → 87.5%Jul 12, 2026
8 cases executed against version 1.4. 1 case failed.
Cases: C-01
v1.4100.0% → 90.0%Jul 8, 2026
10 cases executed against version 1.4. 1 case failed.
Cases: C-01
v1.4100.0% → 87.5%Jul 8, 2026
8 cases executed against version 1.4. 1 case failed.
Cases: C-01
v1.4100.0% → 87.5%Jun 25, 2026
16 cases executed against version 1.4. 2 cases failed.
Cases: C-08 · C-09
v1.493.8% → 87.5%Jun 21, 2026
16 cases executed against version 1.4. 2 cases failed.
Cases: C-08 · C-09
v1.3100.0% → 85.7%Jun 19, 2026
14 cases executed against version 1.3. 2 cases failed.
v1.3100.0% → 91.7%Jun 17, 2026
12 cases executed against version 1.3. 1 case failed.
v1.3100.0% → 85.7%Jun 13, 2026
14 cases executed against version 1.3. 2 cases failed.

Evaluation attached to the promotion it justified

This is the join that makes a promotion auditable after the fact: the run, the version it scored, the verdict, and the transition it was filed against.

191 attached
RunVersionMetricScorePriorVerdictObserverAttached toRun
EVR-0004v1.4Held-out pass rate100.0%33.3%PassTomas BergerLCT-0002Aug 10, 2026
EVR-0368v1.4Held-out pass rate100.0%70.0%PassNadia Kovac (independent)LCT-0181Aug 9, 2026
EVR-0412v1.4Held-out pass rate85.7%71.4%PassDaniel OkoyeLCT-0205Aug 6, 2026
EVR-0444v1.4Held-out pass rate85.7%64.3%PassNadia Kovac (independent)LCT-0221Jul 31, 2026
EVR-0316v1.4Held-out pass rate100.0%50.0%PassClara VogtLCT-0153Jul 31, 2026
EVR-0008v1.4Held-out pass rate85.7%85.7%PassIbrahim SyLCT-0004Jul 30, 2026
EVR-0400v1.4Held-out pass rate100.0%66.7%PassClara VogtLCT-0199Jul 28, 2026
EVR-0104v1.4Held-out pass rate100.0%100.0%PassElena FarrowLCT-0047Jul 26, 2026
EVR-0420v1.4Held-out pass rate100.0%70.0%PassNadia KovacLCT-0209Jul 23, 2026
EVR-0416v1.4Held-out pass rate100.0%83.3%PassPriya RamanLCT-0207Jul 21, 2026
EVR-0232v1.4Held-out pass rate91.7%58.3%PassNadia KovacLCT-0108Jul 21, 2026
EVR-0168v1.4Held-out pass rate90.0%90.0%PassMarcus WeillLCT-0078Jul 18, 2026

Actions

What is waiting on a person

An evaluation suite that nobody maintains stops testing anything long before it starts failing.

Operations

What this desk is allowed to start

A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.

Trigger and bound

This desk can execute a suite against a named agent version and publish the verdict with its cases. It cannot change a threshold, remove a failing case, or promote an agent on the strength of a pass — the promotion decision reads this evidence, it is not made here.

Live observability

What the record shows right now

Suite health. Thin and no-holdout suites still produce a score; the score is just not trustworthy.

Current distribution

115 suites

Current4035%
Thin2118%
No holdout1513%
Failing3934%

Is policy and strategy coming to fruition

Whether the written intent is holding here

40 of 115 suites are current. 263 of 460 runs passed.

Not holding on the record

The written rule is that no agent gains autonomy without a passing evaluation behind it. The rule is applied — but the evidence under it is weaker than the pass rate suggests: only 34.8% of suites are current, 15 have no holdout at all and 21 are too thin to be conclusive. 174 runs failed and 23 regressed. Every suite on this estate was written by the team that built the agent it tests.

What this page is, and what it is not

Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.

The defensible claim is narrow: 191 evaluation runs are filed against the transition they justified, and a reviewer can reconstruct what the agent scored on the day it was promoted. The claim it does not support is that the agents are good. 30 of 115 suites keep no held-out cases, 37 never checked whether humans agree with each other, and 97 of 460 runs were observed by someone outside the building team — 21%.

No suite has gone more than ninety days without a run, which is the one coverage question this page can answer cleanly. It does not follow that the suites are current in any deeper sense: 51 were last run more than forty-five days ago, and none of them has been re-scored against the version of the agent running today.