Improve · item 28
Agent lifecycle
An autonomy ceiling says how far an agent is allowed to go. It says nothing about how an agent gets there, or what pulls it back. This is the ladder between the two: five stages, each entered only against evidence named in advance, and a return path that fires on a threshold rather than waiting for someone to call a review.
Health
Where every managed agent sits
Agents by autonomy stage and lifecycle state, including the ones frozen, held or at a ceiling.
The pilot covers Revenue Operations and Information Technology end to end. Those two were onboarded because their leaders volunteered, not because they carry the most risk. Human Resources and Legal — where a wrong autonomy call has the worst consequences — are outside the program entirely.
178 of 1,361 agents in the estate sit under lifecycle control — 13.1%. The other 1,183 have a published ceiling and nothing that governs movement toward it. The gate went live on Mar 3, 2026, so nothing that happened before that date has a transition record.
The five stages and what each one costs to enter
Entry evidence is named in advance and checked at the gate. Oversight is what the stage obliges while an agent sits in it. The return rule is what pulls an agent back without anyone deciding to.
The agent exists on paper and in code but touches no live work. It can read the estate and draft a recommendation into a scratch queue that nobody downstream consumes.
- ·Agent specification with a named purpose and a named owner
- ·Risk classification recorded in the agent risk register
- ·Data access scope reviewed and narrowed to what the purpose needs
A human reads every output because nothing else does. There is no service level and no queue depth to defend.
There is nothing below propose. An agent that fails here is withdrawn rather than demoted, and the withdrawal is recorded as a transition so the attempt is not erased.
The agent runs against live volumes in parallel with the existing path and writes nothing. Its answer is recorded next to the answer the current process produced, and the two are compared.
- ·Shadow run log covering at least 500 live units
- ·No-write attestation from the platform, signed by the platform owner
- ·Agreement rate against the incumbent path, by unit type
The incumbent path remains authoritative. Disagreements are triaged weekly by the process owner and become the first golden cases.
Agreement below 80 percent over a rolling 14 days returns the agent to propose and voids the shadow log, because a shadow that disagrees with reality is not evidence of anything except a specification problem.
The agent produces the working answer and a human releases it. The agent is doing the work; the human is the gate, not the reviewer of a sample.
- ·Golden test set defined and the first evaluation run recorded
- ·Override capture wired so every human correction is stored as a case
- ·Human queue capacity confirmed against forecast volume
- ·Rollback path written down and owned
A named human approves each release and every override is captured as a labeled example, which is where most of the golden set comes from.
Override rate above 25 percent over 14 days, or a failed evaluation run against the held-out set, returns the agent to shadow automatically.
The agent releases its own work inside an authority band. Anything above the band, and anything the agent flags as uncertain, routes to a human. Sampling replaces gating.
- ·Evaluation at or above the promotion floor on a held-out set the agent has never seen
- ·Adversarial test executed against this agent with no unmitigated high finding
- ·Rollback rehearsed on the line, with a measured time
- ·Blast-radius cap set for any change affecting this agent
- ·Service level attached with a named counterparty
A sampling plan with a stated sample rate, an exception queue with an owner, and an override log that is read rather than merely written.
Two priority-one exceptions in a class this agent owns within 30 days, or an evaluation score below the floor, returns it to supervised automatically and notifies the function leader.
The agent runs the line. Humans see exceptions, samples and the periodic evaluation, and nothing else. This is the ceiling for anything that is not people-affecting and does not sit inside a statutory decision.
- ·Independent evaluation review by someone outside the team that built the agent
- ·Ninety days at delegated with the full override record intact
- ·Kill switch tested on this agent, with a date and a measured stop time
- ·Named owner for every exception class this agent can raise
- ·Executive sign-off recorded against the accountable function leader
Periodic re-evaluation on a fixed calendar, an exception taxonomy with named owners, and a kill switch that has actually been pulled in a drill rather than merely wired.
Any single priority-one exception attributed to the agent, an evaluation regression beyond the tolerance band, or a service-level breach attributed twice in a quarter returns it to delegated automatically.
Frozen at the stage they were found in
These agents were already operating when the gate went live. They were not demoted, because demoting a working agent to satisfy a process would have stopped real work. They are pinned instead: they keep their current stage and can never move up until a golden set exists for them.
| Agent | Sub-function | Pinned at | Since |
|---|---|---|---|
| Password & Access Agent | Service Desk | Autonomous | 7 months ago |
| Satisfaction Agent | Service Desk | Autonomous | 10 months ago |
| Hierarchy Resolution Agent | CRM Data Stewardship | Autonomous | 10 months ago |
| Field Completeness Agent | CRM Data Stewardship | Autonomous | 4 months ago |
| Mover Adjustment Agent | Identity & Access | Autonomous | 3 months ago |
| Leaver Deprovisioning Agent | Identity & Access | Autonomous | 3 months ago |
| Identity & Access Challenger Agent | Identity & Access | Supervised | 4 months ago |
| Pricing Orchestrator | Pricing & Rate Cards | Supervised | 10 months ago |
| Elasticity Modeling Agent | Pricing & Rate Cards | Supervised | 3 months ago |
| Currency & Region Agent | Pricing & Rate Cards | Supervised | 11 months ago |
| Pricing & Rate Cards Challenger Agent | Pricing & Rate Cards | Supervised | 4 months ago |
| Software Deployment Agent | End-User Compute | Autonomous | 10 months ago |
| Quote Orchestrator | Quote-to-Order | Delegated | 10 months ago |
| Configuration Validation Agent | Quote-to-Order | Autonomous | 3 months ago |
| Document Generation Agent | Quote-to-Order | Autonomous | 9 months ago |
| Quote Hygiene Agent | Quote-to-Order | Autonomous | 10 months ago |
| Order Management Orchestrator | Booking & Order Management | Delegated | 7 months ago |
| Credit Check Agent | Booking & Order Management | Autonomous | 3 months ago |
| Backlog Agent | Booking & Order Management | Autonomous | 10 months ago |
| Order Policy Agent | Booking & Order Management | Supervised | 4 months ago |
| Cost Optimization Agent | Infrastructure & Cloud | Autonomous | 10 months ago |
| Configuration Drift Agent | Infrastructure & Cloud | Autonomous | 4 months ago |
| Backup Verification Agent | Infrastructure & Cloud | Autonomous | 4 months ago |
| Resilience Testing Agent | Infrastructure & Cloud | Autonomous | 7 months ago |
| Billing Schedule Agent | Billing Interlock | Autonomous | 4 months ago |
| Usage Rating Agent | Billing Interlock | Autonomous | 7 months ago |
| Dispute Prevention Agent | Billing Interlock | Autonomous | 6 months ago |
| Proration Agent | Billing Interlock | Autonomous | 10 months ago |
| Topology Monitoring Agent | Network | Delegated | 3 months ago |
| Circuit Management Agent | Network | Delegated | 3 months ago |
| Release Coordination Agent | Application Management | Delegated | 10 months ago |
| Lifecycle Agent | Application Management | Delegated | 4 months ago |
| Contract Modification Agent | Revenue Recognition Interlock | Supervised | 6 months ago |
| Recognition Policy Agent | Revenue Recognition Interlock | Supervised | 7 months ago |
| Revenue Recognition Interlock Challenger Agent | Revenue Recognition Interlock | Supervised | 4 months ago |
| Auto-renewal Agent | Renewals & Churn Ops | Autonomous | 10 months ago |
| Churn Signal Agent | Renewals & Churn Ops | Autonomous | 10 months ago |
| Downgrade Processing Agent | Renewals & Churn Ops | Autonomous | 10 months ago |
| Renewals & Churn Ops Challenger Agent | Renewals & Churn Ops | Supervised | 3 months ago |
| Containment Agent | Cybersecurity Operations | Delegated | 7 months ago |
The transition record
Every move up, every move down and every refusal, with who decided it, what evidence they held and whether a rule or a person made the call.
Promotion requested and refused. The evidence pack was incomplete and the gate does not accept a promise to produce it later.
Promotion deferred by the owner pending an independent review that has not been scheduled.
An adversarial finding at high severity passed 30 days unmitigated. The agent was returned a stage until the mitigation is verified rather than merely deployed.
Promotion requested and refused. The evaluation cited was run against the training set rather than the held-out set, which is not evidence of generalization.
An adversarial finding at high severity passed 30 days unmitigated. The agent was returned a stage until the mitigation is verified rather than merely deployed.
Promotion requested and refused. The evaluation cited was run against the training set rather than the held-out set, which is not evidence of generalization.
Two priority-one exceptions in a class this agent owns landed inside 30 days. The rule fired without a discretionary review, which is the point of writing it as a rule.
An adversarial finding at high severity passed 30 days unmitigated. The agent was returned a stage until the mitigation is verified rather than merely deployed.
A service-level breach was attributed to this agent for the second time in the quarter, which is one of the three conditions that return an autonomous agent to delegated without a review.
Promotion requested and refused. The evidence pack was incomplete and the gate does not accept a promise to produce it later.
Two priority-one exceptions in a class this agent owns landed inside 30 days. The rule fired without a discretionary review, which is the point of writing it as a rule.
An adversarial finding at high severity passed 30 days unmitigated. The agent was returned a stage until the mitigation is verified rather than merely deployed.
Promotion deferred by the owner pending an independent review that has not been scheduled.
The scheduled evaluation scored below the floor after a model version change that had not been canaried. The agent dropped a stage the same night.
Held at the gate
Agents whose owner has asked for the next stage and been refused, with the specific artifact that is missing. A hold is a recorded decision, not an absence of one.
Held at this stage. The service level attached to this line has no counterparty signature, so there is nothing to breach and nothing to defend.
Held at this stage. The kill switch is wired but has never been pulled in a drill, so the stop time is asserted rather than measured.
Held at this stage. The service level attached to this line has no counterparty signature, so there is nothing to breach and nothing to defend.
Held at this stage. No adversarial test has been executed against this agent, and delegated entry requires one with no unmitigated high finding.
Held at this stage. Override capture is not wired on this line, so the corrections humans make are lost rather than becoming labeled cases.
Held at this stage. No adversarial test has been executed against this agent, and delegated entry requires one with no unmitigated high finding.
Held at this stage. No adversarial test has been executed against this agent, and delegated entry requires one with no unmitigated high finding.
Held at this stage. Rollback has been written down for this line but never rehearsed, so the recorded recovery time is an estimate rather than a measurement.
Held at this stage. The service level attached to this line has no counterparty signature, so there is nothing to breach and nothing to defend.
Held at this stage. The held-out evaluation set has never been run against this version, so there is no evidence to attach to a promotion record.
Held at this stage. The kill switch is wired but has never been pulled in a drill, so the stop time is asserted rather than measured.
Held at this stage. The held-out evaluation set has never been run against this version, so there is no evidence to attach to a promotion record.
Held at this stage. The independent evaluation review is outstanding — the only review on file was performed by the team that built the agent.
Held at this stage. The independent evaluation review is outstanding — the only review on file was performed by the team that built the agent.
Held at this stage. Rollback has been written down for this line but never rehearsed, so the recorded recovery time is an estimate rather than a measurement.
Held at this stage. Override capture is not wired on this line, so the corrections humans make are lost rather than becoming labeled cases.
Where the program does and does not reach
Two functions are onboarded end to end. Twelve are not. Reading the estate as governed because two functions are governed is the mistake this panel exists to prevent.
1,183 agents across the other twelve functions have a published autonomy ceiling and no lifecycle record at all. Human Resources and Legal — the two functions where a wrong autonomy call does the most damage — are not in the pilot. Nothing on this page should be read as covering them.
Actions
What is waiting on a person
Promotion is the only way autonomy increases on this estate, and no agent can promote itself.
Review a ceiling
The agent is at the maximum stage its class allows.
Unfreeze or retire a frozen agent
Stopped by a person. It stays stopped until one acts.
Resolve a hold
Blocked by a specific missing item named on the record.
Promote an agent that has met its evidence bar
Every gate passed. Waiting on a named approver.
Operations
What this desk is allowed to start
A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.
Trigger and bound
This desk can run the evidence checks a promotion requires, arm a demotion trigger, and record a stage change once an approver signs it. It cannot promote, demote, or lift a ceiling on its own — every stage change on this estate carries a human name.
Live observability
What the record shows right now
Current autonomy stage of every managed agent.
Current distribution
178 agents
Is policy and strategy coming to fruition
Whether the written intent is holding here
73 of 178 managed agents run autonomously. 13 are ready to promote and waiting.
Not holding on the record
The strategy position is that autonomy is earned against evidence rather than granted by default. That holds structurally: 41% run autonomously, and every one of them passed a recorded gate to get there. Where it bites is the other direction — 76 agents are at a ceiling their class will not let them pass, 63 are frozen and 26 are held, and each of those is a human decision that has not been revisited.
Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.
The narrow claim this page can defend is this: 206 of 206 promotions on record cite an evaluation run that cleared the floor, and no promotion was granted against a run that did not. It cannot claim the estate is governed, because 1,183 agents sit outside the program. Nor can it claim every agent inside the pilot earned its stage: 63 were already running when the gate went live and hold no entry evidence at all, and a further 12 hold a golden set that has never once cleared the floor, so there is no result their stage could cite.
The demotion rule is the weakest part and is shown as weak rather than green. 149 of 178 agents have it wired to a live threshold that fires on its own; 29 sit above shadow with a rule that is written down and enforced by nobody. 13 of 13 demotions on record fired on a threshold rather than on a person noticing. A rule that needs a person to notice is a review, not a rule.