Information Technology · identity-and-access Identity & Access
Identity & Access Challenger Agent
challenger agent in the strategy layer, owned by Marcus Weill (Director, Model Governance). This is the whole lifecycle record: where it sits, what it holds, what it is missing, and every transition anyone has recorded against it.
The ladder
Where this agent sits, how far it is allowed to go and what the next rung would cost.
Nothing on file. This agent was already running when the gate went live on Mar 3, 2026.
- ·Golden test set defined and the first evaluation run recorded
- ·Override capture wired so every human correction is stored as a case
- ·Human queue capacity confirmed against forecast volume
- ·Rollback path written down and owned
This agent was already running at this stage when the lifecycle gate was introduced. Its entry evidence was never produced. It is frozen at its current stage — it cannot be promoted, and the missing evidence is listed against it — but it was not demoted, because demoting a working line to satisfy a records gap would have been theatre.
Override rate above 25 percent over 14 days, or a failed evaluation run against the held-out set, returns the agent to shadow automatically.
Wired to a live threshold. It fires and moves the agent down without anyone deciding to.
Running at a stage it reached before the gate existed. The platform holds no promotion record for it.
Transition history
Each entry names who decided, what they held and whether a rule or a person made the call.
No transitions recorded. This agent predates the gate, which went live on Mar 3, 2026. It reached its current stage before anyone was keeping a record, and that absence is why it is frozen rather than shown as compliant.
Evaluation record
The golden set this agent is measured against, and every run scored against it.
Nothing has been written that this agent can be tested against. Without a golden set there is no evidence a promotion could cite, which is exactly why this agent cannot move up regardless of how well it appears to be performing.
Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.
This agent has no golden set, so there is nothing here to score. The record shows an absence rather than a pass, which is the honest reading: an untested agent is not a safe agent that has not been checked yet, it is an agent whose behavior nobody has established.