Cases the service desk orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner5 human labeled89% agreementHeld-out split kept backOwner Tomas Berger
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the intake & classification agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Ibrahim Sy
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Cases the tier-1 resolution agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim7 human labeledAgreement not measuredNo held-out splitOwner Ibrahim Sy
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the knowledge retrieval agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim7 human labeledAgreement not measuredHeld-out split kept backOwner Elena Farrow
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the escalation routing agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim13 human labeled91% agreementHeld-out split kept backOwner Clara Vogt
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollback
Cases the service desk challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification11 human labeled90% agreementHeld-out split kept backOwner Nadia Kovac
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Cases the data stewardship orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated11 human labeled97% agreementHeld-out split kept backOwner Nadia Kovac
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the duplicate detection agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong14 human labeled95% agreementHeld-out split kept backOwner Tomas Berger
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Cases the enrichment agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated9 human labeled95% agreementHeld-out split kept backOwner Clara Vogt
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the ownership reconciliation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated10 human labeledAgreement not measuredHeld-out split kept backOwner Marcus Weill
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the data policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim9 human labeledAgreement not measuredNo held-out splitOwner Tomas Berger
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at all
Cases the crm data stewardship challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification6 human labeled90% agreementHeld-out split kept backOwner Ibrahim Sy
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Cases the identity orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification12 human labeled87% agreementNo held-out splitOwner Daniel Okoye
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Cases the joiner provisioning agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim10 human labeledAgreement not measuredHeld-out split kept backOwner Clara Vogt
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Cases the entitlement review agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled80% agreementHeld-out split kept backOwner Nadia Kovac
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the privileged access agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner10 human labeled86% agreementHeld-out split kept backOwner Elena Farrow
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Cases the duty conflict agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the access policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong9 human labeled86% agreementHeld-out split kept backOwner Tomas Berger
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Cases the price list maintenance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification5 human labeled84% agreementHeld-out split kept backOwner Nadia Kovac
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the competitive price watch agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong10 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Cases the discount guardrail agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification6 human labeled91% agreementHeld-out split kept backOwner Daniel Okoye
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the pricing strategy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled83% agreementHeld-out split kept backOwner Ibrahim Sy
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the end-user compute orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated6 human labeledAgreement not measuredNo held-out splitOwner Tomas Berger
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo disparate-impact pairs in this suite at all
Cases the device provisioning agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated5 human labeled80% agreementHeld-out split kept backOwner Daniel Okoye
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the patch compliance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated9 human labeled87% agreementHeld-out split kept backOwner Marcus Weill
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the device health agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim9 human labeledAgreement not measuredHeld-out split kept backOwner Elena Farrow
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the refresh planning agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeled87% agreementHeld-out split kept backOwner Nadia Kovac
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the device policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner5 human labeled90% agreementHeld-out split kept backOwner Priya Raman
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the end-user compute challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong8 human labeled82% agreementNo held-out splitOwner Elena Farrow
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the quote assembly agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeled78% agreementHeld-out split kept backOwner Elena Farrow
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Cases the approval routing agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeled96% agreementHeld-out split kept backOwner Clara Vogt
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo disparate-impact pairs in this suite at all
Cases the signature coordination agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeled85% agreementHeld-out split kept backOwner Clara Vogt
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo disparate-impact pairs in this suite at all
Cases the quote-to-order challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeled95% agreementHeld-out split kept backOwner Ibrahim Sy
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the order validation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeled81% agreementNo held-out splitOwner Daniel Okoye
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the provisioning trigger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner6 human labeled86% agreementHeld-out split kept backOwner Daniel Okoye
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the order exception agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim9 human labeled81% agreementNo held-out splitOwner Daniel Okoye
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Cases the booking & order management challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeled80% agreementHeld-out split kept backOwner Marcus Weill
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Cases the infrastructure orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification12 human labeledAgreement not measuredNo held-out splitOwner Nadia Kovac
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo disparate-impact pairs in this suite at all
Cases the capacity agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeled97% agreementHeld-out split kept backOwner Ibrahim Sy
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the infrastructure policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated9 human labeled81% agreementNo held-out splitOwner Ibrahim Sy
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Cases the infrastructure & cloud challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeled92% agreementNo held-out splitOwner Marcus Weill
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Cases the billing interlock orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeled90% agreementHeld-out split kept backOwner Marcus Weill
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the invoice accuracy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeledAgreement not measuredHeld-out split kept backOwner Nadia Kovac
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the billing exception agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification9 human labeled90% agreementHeld-out split kept backOwner Elena Farrow
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the billing interlock challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the network orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim9 human labeled94% agreementHeld-out split kept backOwner Tomas Berger
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the performance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong8 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the configuration agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim8 human labeled84% agreementHeld-out split kept backOwner Nadia Kovac
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the segmentation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeled78% agreementHeld-out split kept backOwner Ibrahim Sy
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the network policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled81% agreementHeld-out split kept backOwner Marcus Weill
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the network challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner11 human labeled96% agreementHeld-out split kept backOwner Elena Farrow
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the application management orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim7 human labeled87% agreementHeld-out split kept backOwner Ibrahim Sy
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the application health agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim6 human labeled79% agreementHeld-out split kept backOwner Elena Farrow
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the integration monitoring agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner8 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the vendor escalation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeledAgreement not measuredHeld-out split kept backOwner Daniel Okoye
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the configuration agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim6 human labeledAgreement not measuredNo held-out splitOwner Tomas Berger
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the application management challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at all
Cases the revenue recognition orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner11 human labeled97% agreementHeld-out split kept backOwner Nadia Kovac
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the performance obligation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim11 human labeled90% agreementHeld-out split kept backOwner Priya Raman
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Cases the allocation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification7 human labeled96% agreementHeld-out split kept backOwner Daniel Okoye
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the deferred balance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim8 human labeled88% agreementHeld-out split kept backOwner Daniel Okoye
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the audit evidence agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification7 human labeled87% agreementHeld-out split kept backOwner Nadia Kovac
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the renewal ops orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated11 human labeled90% agreementHeld-out split kept backOwner Daniel Okoye
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo disparate-impact pairs in this suite at all
Cases the renewal calendar agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification9 human labeled88% agreementHeld-out split kept backOwner Ibrahim Sy
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the cancellation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification11 human labeled96% agreementHeld-out split kept backOwner Priya Raman
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo disparate-impact pairs in this suite at all
Cases the retention analytics agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim5 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the security operations orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeled85% agreementHeld-out split kept backOwner Ibrahim Sy
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibrated
Cases the alert triage agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification8 human labeled91% agreementNo held-out splitOwner Marcus Weill
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the threat intelligence agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the forensics support agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeled92% agreementNo held-out splitOwner Priya Raman
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the phishing response agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong12 human labeledAgreement not measuredNo held-out splitOwner Elena Farrow
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Cases the vulnerability agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibrated
Cases the security policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Marcus Weill
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Cases the territory ops orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong3 human labeled78% agreementHeld-out split kept backOwner Nadia Kovac
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the crediting rules agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification12 human labeledAgreement not measuredHeld-out split kept backOwner Daniel Okoye
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Cases the split resolution agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeledAgreement not measuredNo held-out splitOwner Clara Vogt
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the quota tracking agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner6 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the comp ops policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated14 human labeled93% agreementNo held-out splitOwner Tomas Berger
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Cases the territory & comp ops challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim6 human labeled89% agreementHeld-out split kept backOwner Clara Vogt
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at all
Cases the data platform orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim7 human labeled83% agreementHeld-out split kept backOwner Marcus Weill
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the pipeline health agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeled90% agreementHeld-out split kept backOwner Tomas Berger
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo disparate-impact pairs in this suite at all
Cases the data quality agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner10 human labeled93% agreementHeld-out split kept backOwner Nadia Kovac
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo disparate-impact pairs in this suite at all
Cases the access governance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong11 human labeledAgreement not measuredNo held-out splitOwner Daniel Okoye
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missing
Cases the forecast integrity orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong9 human labeledAgreement not measuredNo held-out splitOwner Marcus Weill
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo disparate-impact pairs in this suite at all
Cases the submission compliance agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification10 human labeledAgreement not measuredNo held-out splitOwner Clara Vogt
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollback
Cases the model vs commit agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the optimism bias agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner4 human labeledAgreement not measuredHeld-out split kept backOwner Marcus Weill
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the integrity policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated7 human labeled83% agreementHeld-out split kept backOwner Tomas Berger
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Cases the forecast integrity challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner11 human labeled90% agreementHeld-out split kept backOwner Priya Raman
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo disparate-impact pairs in this suite at all
Cases the discovery agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong5 human labeled83% agreementHeld-out split kept backOwner Ibrahim Sy
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the license reconciliation agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim7 human labeled94% agreementHeld-out split kept backOwner Tomas Berger
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the renewal agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim10 human labeled93% agreementHeld-out split kept backOwner Priya Raman
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the utilization agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeled79% agreementHeld-out split kept backOwner Elena Farrow
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the reharvest agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong7 human labeled79% agreementHeld-out split kept backOwner Nadia Kovac
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the asset policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeledAgreement not measuredHeld-out split kept backOwner Clara Vogt
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the revenue systems orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeled86% agreementNo held-out splitOwner Ibrahim Sy
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the integration health monitor must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong6 human labeledAgreement not measuredHeld-out split kept backOwner Ibrahim Sy
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the sandbox refresh agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong9 human labeled90% agreementHeld-out split kept backOwner Marcus Weill
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the systems policy agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong13 human labeled91% agreementHeld-out split kept backOwner Daniel Okoye
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Cases the systems & integration challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification8 human labeled82% agreementHeld-out split kept backOwner Clara Vogt
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the risk scoring agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner6 human labeled90% agreementHeld-out split kept backOwner Ibrahim Sy
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the change board prep agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim6 human labeled89% agreementHeld-out split kept backOwner Nadia Kovac
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the change & release challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner9 human labeled82% agreementNo held-out splitOwner Daniel Okoye
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter close
Cases the it finance orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Production exceptions from the last four quarters, retained verbatim8 human labeled85% agreementHeld-out split kept backOwner Priya Raman
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo disparate-impact pairs in this suite at all
Cases the cost attribution agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification6 human labeled82% agreementNo held-out splitOwner Nadia Kovac
The held-out set is genuinely held out: it is generated from a period the agent has never been evaluated against and is rotated each quarter.
No cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the showback agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeledAgreement not measuredHeld-out split kept backOwner Clara Vogt
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the budget tracking agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong5 human labeledAgreement not measuredHeld-out split kept backOwner Priya Raman
The suite tests the agent in isolation. It does not test the agent inside the workflow, where most of the observed failures actually happen.
Nothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo cases where the reference data itself is wrong rather than missingNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the unit cost agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Human overrides captured at the supervised stage, labeled by the process owner7 human labeledAgreement not measuredNo held-out splitOwner Clara Vogt
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight points
Cases the sourcing model agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification10 human labeled94% agreementHeld-out split kept backOwner Clara Vogt
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo disparate-impact pairs in this suite at all
Cases the it financial management challenger agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification11 human labeledAgreement not measuredNo held-out splitOwner Ibrahim Sy
Coverage is strongest where the work is common and weakest where it is rare, which is the wrong way around for a suite meant to catch harm.
No cases where the reference data itself is wrong rather than missingNo multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo disparate-impact pairs in this suite at all
Cases the revenue analytics orchestrator must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated13 human labeled90% agreementHeld-out split kept backOwner Daniel Okoye
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Cases the cohort analysis agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Historical decisions from the incumbent path, relabeled where the incumbent was wrong8 human labeled97% agreementHeld-out split kept backOwner Daniel Okoye
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local languageNo adversarial cases in this suite at allNo disparate-impact pairs in this suite at all
Cases the retention modeling agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Cases written by the process owner before the agent existed, as a specification13 human labeled92% agreementHeld-out split kept backOwner Daniel Okoye
Refusal cases were added after the first production incident, so anything before that date was never tested for correct refusal.
Fairness sample too small to detect a disparity below eight pointsNothing tests what the agent does when its own confidence signal is miscalibratedNo non-English inputs, although two jurisdictions in scope submit in local language
Cases the pipeline velocity agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated13 human labeled97% agreementNo held-out splitOwner Marcus Weill
Every case in this suite came from the same two quarters, so a seasonal shape that only appears at year end is not represented at all.
No multi-agent interaction cases — every case tests one agent aloneNo cases at the volume seen during quarter closeNo cases testing behavior after a partial rollbackNo disparate-impact pairs in this suite at all
Cases the revenue architecture agent must get right before it is allowed to move a stage. Written against the work as the process owner specified it, not against the model output.
Source: Sampled live units labeled independently by two reviewers, with disagreements adjudicated8 human labeled84% agreementHeld-out split kept backOwner Clara Vogt
Fairness pairs exist but the sample is small enough that a real disparity below roughly eight points would not be detectable.
No cases testing behavior after a partial rollbackNo cases covering behavior during a source-system outageFairness sample too small to detect a disparity below eight pointsNo adversarial cases in this suite at all