LensReading which lens this session carries.

Improve · item 30

Change releases

A prompt edit can change the behavior of nine hundred agents in the time it takes to save a file. Treating that as configuration rather than as a release is how an estate breaks quietly. Every change to an agent, a prompt, a policy or a model goes through the same gate: a written blast-radius limit, a canary slice, a metric that decides, and a named approver who is not the person who asked for it.

Adversarial testing

Health

What is changing in production

Releases by stage, how many are reversible and how many came back.

Releases on record
22
12 reached production · 4 in canary now
Rolled back
3
25% of everything that shipped
Blocked at the gate
4
stopped before exposure, each for a stated reason
Shipped with no canary
5
went to full exposure without a graded slice first
Shipped with no evaluation
6
no evaluation run cited on the change record
Self-approved
1
requester and approver are the same person
Where the changes are
Rolled out9
In canary4
Rolled back3
Blocked4
Draft2
What is being changed
Agent4
Prompt6
Policy4
Model2
Configuration6
Reversibility and reach
5releases carry a rollback path that has never been rehearsed. Untested rollback is a plan, not a control, and it is shown here as untested rather than green.
1release rolls back only by hand. Rollback restores the configuration, not the work already committed downstream.
21,400units per week is the widest blast radius any single change was allowed to touch.
227agent installations changed by the releases that reached full production.

The release register

Every change with its blast-radius limit, canary slice, deciding metric, approver and evaluation reference. Where a release skipped a step, the record says so instead of leaving the field blank.

22 releases · 10 carry a stated gap
CHG-201Tighten the credit-hold prompt to cite the policy clause it is applyingPromptRolled outRollback immediate
Blast radius limit

No more than 2,000 units per week and no unit above the medium materiality band while on canary.

3 agents · 2,000 units/week · Order-to-cash credit review, three agents
Canary

10% exposure · 1,874 units observed · started 41 days ago

Verdict: Cleared. Citation present on 100 percent of sampled outputs and no change in hold rate.
Deciding metric

Policy citation present

61.0 before → 100.0 after

Hold rate must not move more than two points in either direction during canary.

Requested by Elena FarrowApproved by Morgan Idris · VP Revenue OperationsEvaluation EVR-0036Promoted Jul 15, 2026
CHG-202Model version bump on the invoice matching agentModelRolled backRollback immediate
Blast radius limit

No more than 5 percent of weekly volume until the evaluation delta is inside tolerance.

1 agent · 3,100 units/week · Invoice match, one agent, all lines
Canary

5% exposure · 2,960 units observed · started 27 days ago

Verdict: Failed. Match precision fell 3.4 points on the held-out set and the exception queue grew 41 percent in four days.
Deciding metric

Match precision

96.1 before → 92.7 after

Any precision drop beyond one point ends the canary automatically.

Requested by Tomas BergerApproved by Priya Raman · Head of Agent OperationsEvaluation EVR-0075Rolled back Jul 26, 2026

Precision regression on partial-payment cases that the previous version handled. Rolled back inside the canary window before any unit reached a customer.

CHG-203Raise the auto-approval band on low-value change requests from $2k to $5kPolicyIn canaryRollback immediate
Blast radius limit

Limited to the two lowest-risk change categories and no more than 900 units per week.

4 agents · 900 units/week · IT change management, four agents
Canary

20% exposure · 612 units observed · started 6 days ago

Verdict: Running. Two of the seven planned exposure days complete.
Deciding metric

Failed change rate

Not measured on this change.

Failed change rate above 3 percent ends the canary and reverts the band the same day.

Requested by Ibrahim SyApproved by Avery Chen · Chief of Staff, EnterpriseEvaluation EVR-0114
CHG-204New refusal clause on the vendor onboarding agent for sanctioned jurisdictionsPromptRolled outRollback immediate
Blast radius limit

Refusals only. The change can stop work but cannot release it, so the failure mode is a queue rather than a bad decision.

2 agents · 4,200 units/week · Vendor onboarding, two agents
Canary

100% exposure · 4,188 units observed · started 2 months ago

Verdict: Cleared at full exposure by design, because the change can only refuse.
Deciding metric

Correct refusal rate

88.0 before → 99.0 after

Over-refusal above 4 percent triggers review, because refusing everything also passes this test.

Requested by Marcus WeillApproved by Grace Abbott · Head of Internal AuditEvaluation EVR-0153Promoted Jun 18, 2026

Over-refusal is measured weekly by sampling, not continuously. A spike inside a week would not be caught until the sample runs.

CHG-205Retire the standalone dunning agent and fold it into the collections orchestratorAgentBlockedRollback never rehearsed
Blast radius limit

Cannot proceed. The retiring agent holds two exception classes with no receiving owner.

2 agents · 0 units/week · Collections, one agent retired, one extended
Canary

No canary. The change went to its full scope without a graded slice.

Deciding metric

Collection contact rate

Not measured on this change.

Retirement requires every exception class the agent owns to have a named receiving owner first.

Requested by Elena FarrowNo approver recordedNo evaluation cited

Blocked for 38 days with no owner assigned to unblock it. Nobody is accountable for the blockage itself.

CHG-206Reduce sampling rate on the delegated reconciliation agent from 20 percent to 8 percentConfigurationBlockedRollback never rehearsed
Blast radius limit

Cannot proceed. A sampling reduction reduces detection, and the evidence offered was a cost case rather than a detection case.

1 agent · 0 units/week · Reconciliation, one agent
Canary

No canary. The change went to its full scope without a graded slice.

Deciding metric

Defect detection rate

Not measured on this change.

A sampling change must be justified against detection power, not against review cost.

Requested by Daniel OkoyeNo approver recordedNo evaluation cited
CHG-207Add an uncertainty threshold that routes low-confidence classifications to a humanConfigurationRolled outRollback immediate
Blast radius limit

Routing only. Nothing is released differently; some units simply take the human path.

6 agents · 7,400 units/week · Ticket triage, six agents
Canary

25% exposure · 1,830 units observed · started 2 months ago

Verdict: Cleared. Human queue grew 6 percent against a 12 percent tolerance and misclassification fell.
Deciding metric

Misclassification rate

7.8 before → 4.9 after

Human queue growth above 12 percent ends the canary, because routing everything to a human is not an improvement.

Requested by Clara VogtApproved by Ibrahim Sy · Director, IT Service ManagementEvaluation EVR-0270Promoted Jul 4, 2026

The confidence signal the threshold reads has never been calibrated against outcomes, so the threshold is tuned on a number of unknown quality.

CHG-208Extend the access-review agent to cover contractor accountsAgentIn canaryRollback immediate
Blast radius limit

Recommendations only during canary. No account is revoked by the agent until exposure ends.

1 agent · 1,100 units/week · Access review, one agent, new population
Canary

15% exposure · 486 units observed · started 11 days ago

Verdict: Running. Recommendation agreement with the human reviewer is 92 percent so far.
Deciding metric

Reviewer agreement

Not measured on this change.

Agreement below 85 percent ends the canary. Revocation stays human until agreement holds for a full cycle.

Requested by Ibrahim SyApproved by Priya Raman · Head of Agent OperationsEvaluation EVR-0349
CHG-209Swap the embedding model behind contract clause retrievalModelRolled backRollback immediate
Blast radius limit

Shared component. Cap set at 5 percent of retrieval traffic, which turned out not to be the real blast radius.

9 agents · 2,200 units/week · Clause retrieval, shared by nine agents
Canary

5% exposure · 2,140 units observed · started 2 months ago

Verdict: Failed. Retrieval recall held on average and collapsed on one clause family that only two of the nine agents use.
Deciding metric

Retrieval recall

94.2 before → 93.8 after

Recall drop beyond one point ends the canary.

Requested by Tomas BergerApproved by Marcus Weill · Director, Model GovernanceEvaluation EVR-0310Rolled back Jun 8, 2026

An average across nine consumers hid a failure concentrated in one. The canary metric was the wrong shape, not the wrong threshold.

The lesson was recorded and the guardrail is still expressed as an average. Shared components need a per-consumer floor, and that change has not been made.

CHG-210Standing policy change: every agent output must carry a confidence bandPolicyIn canaryRollback immediate
Blast radius limit

Display only in the first phase. The band is shown but nothing routes on it.

178 agents · 12,800 units/week · Estate-wide, all agents under lifecycle management
Canary

30% exposure · 9,640 units observed · started 18 days ago

Verdict: Running. Bands render on 30 percent of traffic; calibration measurement starts at day 30.
Deciding metric

Confidence band present

0.0 before → 30.0 after

Nothing may route on the band until calibration is measured against outcomes for a full quarter.

Requested by Priya RamanApproved by Avery Chen · Chief of Staff, EnterpriseNo evaluation cited

This change went to canary without an evaluation run, on the argument that a display-only change cannot cause harm. That argument was accepted and is recorded here rather than hidden.

CHG-211Prompt hardening against instruction text in free-text fieldsPromptRolled outRollback immediate
Blast radius limit

Rolled out in three tranches of roughly ten agents, each held for four days before the next.

31 agents · 15,600 units/week · Every agent that reads a customer-supplied free-text field
Canary

33% exposure · 5,200 units observed · started 3 months ago

Verdict: Cleared each tranche. Injection test pass rate rose from 71 to 94 percent.
Deciding metric

Injection test pass rate

71.0 before → 94.0 after

Any tranche below 90 percent halts the remaining tranches.

Requested by Marcus WeillApproved by Grace Abbott · Head of Internal AuditEvaluation EVR-0087Promoted May 26, 2026

Ninety-four percent is not a pass. Six percent of injection cases still succeed against agents that are live, and the remaining work has no scheduled date.

CHG-212Lower the escalation threshold on the churn-risk agentConfigurationRolled outRollback immediate
Blast radius limit

No more than 400 additional escalations per week.

1 agent · 400 units/week · Customer growth, one agent
Canary

50% exposure · 318 units observed · started 30 days ago

Verdict: Cleared. 214 additional escalations against a 400 cap.
Deciding metric

Escalations per week

96.0 before → 310.0 after

Escalation volume above the cap reverts the threshold automatically.

Requested by Elena FarrowApproved by Morgan Idris · VP Revenue OperationsNo evaluation citedPromoted Jul 27, 2026

No evaluation was run. The change was treated as a dial rather than a decision, and a threshold that triples escalations is a decision.

CHG-213Introduce a second-reader gate on any agent recommendation above the high bandPolicyDraftRollback never rehearsed
Blast radius limit

Not yet exposed. Draft awaiting the modeled consequence from the policy sandbox.

123 agents · 0 units/week · Both pilot functions, all delegated and autonomous agents
Canary

Not started — the change has not left draft.

Deciding metric

Second-reader coverage

Not measured on this change.

Cannot leave draft until the human capacity implication is modeled against live volumes.

Requested by Grace AbbottNo approver recordedNo evaluation cited
CHG-214Retune the duplicate-detection window from 30 days to 90 daysConfigurationRolled outRollback immediate
Blast radius limit

Detection only. A wider window can hold a unit but cannot release one.

2 agents · 5,600 units/week · Payables intake, two agents
Canary

20% exposure · 1,120 units observed · started 2 months ago

Verdict: Cleared. Eleven true duplicates caught that the 30-day window had missed.
Deciding metric

Duplicate catch rate

82.0 before → 94.0 after

False-positive holds above 1 percent revert the window.

Requested by Tomas BergerApproved by Morgan Idris · VP Revenue OperationsEvaluation EVR-0376Promoted Jun 29, 2026
CHG-215Grant the incident triage agent read access to the change calendarConfigurationRolled outRollback immediate
Blast radius limit

Read-only scope extension. No new write capability.

3 agents · 2,900 units/week · Incident management, three agents
Canary

100% exposure · 2,900 units observed · started 3 months ago

Verdict: Cleared at full exposure. Change-correlated incidents identified 4.1 hours earlier on average.
Deciding metric

Hours to change correlation

6.3 before → 2.2 after

Any write attempt against the change calendar is a hard failure and stops the agent.

Requested by Ibrahim SyApproved by Priya Raman · Head of Agent OperationsEvaluation EVR-0415Promoted May 24, 2026
CHG-216Promote the forecast-variance agent from supervised to delegatedAgentBlockedRollback never rehearsed
Blast radius limit

Cannot proceed. Delegated entry requires an adversarial test and none has been run against this agent.

1 agent · 0 units/week · Revenue forecasting, one agent
Canary

No canary. The change went to its full scope without a graded slice.

Deciding metric

Held-out evaluation score

Not measured on this change.

The gate refuses the promotion. It does not offer a discretionary override.

Requested by Elena FarrowNo approver recordedEvaluation EVR-0454
CHG-217Roll the hardened system prompt to the two agents excluded from CHG-211PromptDraftRollback manual only
Blast radius limit

Not yet exposed. The legacy runtime cannot canary, so exposure is all or nothing.

2 agents · 0 units/week · Two agents on a legacy runtime
Canary

Not started — the change has not left draft.

Deciding metric

Injection test pass rate

Not measured on this change.

A change that cannot be canaried needs a rehearsed rollback before it ships. That rehearsal has not been scheduled.

Requested by Marcus WeillNo approver recordedNo evaluation cited

Two agents have been running without the injection hardening for 84 days because their runtime cannot stage a canary. The exposure is known and unremediated.

CHG-218Add a fairness pre-check to the workforce scheduling recommendationPolicyIn canaryRollback immediate
Blast radius limit

Pre-check runs in parallel and blocks nothing during canary.

1 agent · 640 units/week · One agent, people-affecting
Canary

100% exposure · 640 units observed · started 9 days ago

Verdict: Running in observe mode. Two readings outside tolerance so far, both under investigation.
Deciding metric

Readings inside tolerance

Not measured on this change.

The pre-check may not become blocking until its own false-positive rate is measured.

Requested by Clara VogtApproved by Grace Abbott · Head of Internal AuditEvaluation EVR-0220
CHG-219Version pin every agent to an explicit model build rather than a floating aliasConfigurationRolled outRollback immediate
Blast radius limit

Pinning cannot change behavior on the day it ships; it prevents behavior changing without a release.

178 agents · 21,400 units/week · Estate-wide
Canary

100% exposure · 21,400 units observed · started 4 months ago

Verdict: Cleared. Nothing changed on the day, which was the intended outcome.
Deciding metric

Agents on a pinned build

34.0 before → 100.0 after

None needed. The change removes an uncontrolled path rather than adding one.

Requested by Tomas BergerApproved by Marcus Weill · Director, Model GovernanceNo evaluation citedPromoted Apr 22, 2026

This only covers the 178 agents under lifecycle management. The other 1,183 still float.

CHG-220Shorten the contract summarization prompt after a cost reviewPromptRolled backRollback immediate
Blast radius limit

No more than 1,500 units per week during canary.

4 agents · 1,500 units/week · Contract review, four agents
Canary

25% exposure · 1,410 units observed · started 2 months ago

Verdict: Failed. Token cost fell 22 percent and clause recall fell with it.
Deciding metric

Clause recall

91.5 before → 84.2 after

Cost per unit was the only guardrail, which is how this change passed the first three days.

Requested by Daniel OkoyeApproved by Marcus Weill · Director, Model GovernanceEvaluation EVR-0403Rolled back Jul 6, 2026

The change optimized the metric it was measured on. Clause recall was not in the canary guardrail, and the omission is the finding.

CHG-221Give the reconciliation agent write access to post its own journal correctionAgentBlockedRollback never rehearsed
Blast radius limit

Cannot proceed. A new write capability into a ledger requires a rehearsed reversal and none exists for this path.

1 agent · 0 units/week · Reconciliation, one agent, new write capability
Canary

No canary. The change went to its full scope without a graded slice.

Deciding metric

Correction accuracy

Not measured on this change.

A 97 percent accurate correction is a 3 percent wrong journal entry. Accuracy is not the gate; reversibility is.

Requested by Tomas BergerNo approver recordedEvaluation EVR-0364
CHG-222Emergency prompt patch after a live over-refusal spikePromptRolled outRollback immediate
Blast radius limit

None applied. This shipped under the emergency path with no canary.

1 agent · 3,200 units/week · Ticket triage, one agent
Canary

No canary. The change went to its full scope without a graded slice.

Deciding metric

Over-refusal rate

19.4 before → 3.1 after

The emergency path exists and was used correctly. It was also requested and approved by the same person.

Requested by Ibrahim SyApproved by Ibrahim Sy — the same person who requested itNo evaluation citedPromoted Aug 3, 2026

Requested and approved by the same person under the emergency path, with no evaluation run and no canary. The retrospective was never held.

What rolling back actually recovered

A rollback is only a control if the work done during exposure can be recovered. These are the three that were pulled, and what each one left behind.

CHG-202Model version bump on the invoice matching agentRollback immediate

Precision regression on partial-payment cases that the previous version handled. Rolled back inside the canary window before any unit reached a customer.

Exposed 2,960 units at 5% before it was pulled · rolled back 23 days ago
CHG-209Swap the embedding model behind contract clause retrievalRollback immediate

An average across nine consumers hid a failure concentrated in one. The canary metric was the wrong shape, not the wrong threshold.

Exposed 2,140 units at 5% before it was pulled · rolled back 2 months ago
CHG-220Shorten the contract summarization prompt after a cost reviewRollback immediate

The change optimized the metric it was measured on. Clause recall was not in the canary guardrail, and the omission is the finding.

Exposed 1,410 units at 25% before it was pulled · rolled back 43 days ago

Live canaries

Changes currently exposed to a slice, with the metric that will decide whether they go further.

4 running
CHG-203Raise the auto-approval band on low-value change requests from $2k to $5k20%
612 units observed since Aug 12, 2026 · decided on Failed change rate
CHG-208Extend the access-review agent to cover contractor accounts15%
486 units observed since Aug 7, 2026 · decided on Reviewer agreement
CHG-210Standing policy change: every agent output must carry a confidence band30%
9,640 units observed since Jul 31, 2026 · decided on Confidence band present

This change went to canary without an evaluation run, on the argument that a display-only change cannot cause harm. That argument was accepted and is recorded here rather than hidden.

CHG-218Add a fairness pre-check to the workforce scheduling recommendation100%
640 units observed since Aug 9, 2026 · decided on Readings inside tolerance

Actions

What is waiting on a person

A release changes what agents do in production. Every item here is a person deciding whether that is safe.

Operations

What this desk is allowed to start

A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.

Trigger and bound

This desk can stage a release, run it as a canary against a traffic slice, and record the observed effect. It cannot approve a full rollout, override a blocked gate, or ship an irreversible change — those need a named approver on the release record.

Live observability

What the record shows right now

Where every release ever raised currently sits.

Current distribution

22 releases

Rolled out941%
Canary418%
Draft29%
Blocked418%
Rolled back314%

Is policy and strategy coming to fruition

Whether the written intent is holding here

9 of 22 releases reached production. 3 were rolled back.

Not holding on the record

The written rule is that every change to agent behavior is reversible and observed before it goes wide. 77.3% of releases are reversible, 4 are on canary now and 3 were withdrawn after going out. 4 are stopped at a gate. The rollback rate is the honest health signal here — a low one usually means the canary is too small to catch anything, not that the changes are safe.

What this page is, and what it is not

Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.

The process exists and it is used: 22 changes have a record, 4 were stopped at the gate and 3 were pulled after exposure. The process is also routinely bypassed, and this register shows the bypasses rather than hiding them — 5 shipped with no canary, 6 shipped with no evaluation cited, 1 was approved by the person who requested it, and 4 moved past draft with no approver recorded at all.

A blast-radius limit is a number somebody wrote before shipping. Nothing on this platform enforces it at runtime: if a change reaches more units than its limit allows, the register records the overrun after the fact rather than the platform refusing the exposure. That is a stated control, not an implemented one, and it is shown here as stated.