Improve · item 30
Change releases
A prompt edit can change the behavior of nine hundred agents in the time it takes to save a file. Treating that as configuration rather than as a release is how an estate breaks quietly. Every change to an agent, a prompt, a policy or a model goes through the same gate: a written blast-radius limit, a canary slice, a metric that decides, and a named approver who is not the person who asked for it.
Health
What is changing in production
Releases by stage, how many are reversible and how many came back.
The release register
Every change with its blast-radius limit, canary slice, deciding metric, approver and evaluation reference. Where a release skipped a step, the record says so instead of leaving the field blank.
No more than 2,000 units per week and no unit above the medium materiality band while on canary.
10% exposure · 1,874 units observed · started 41 days ago
Policy citation present
Hold rate must not move more than two points in either direction during canary.
No more than 5 percent of weekly volume until the evaluation delta is inside tolerance.
5% exposure · 2,960 units observed · started 27 days ago
Match precision
Any precision drop beyond one point ends the canary automatically.
Precision regression on partial-payment cases that the previous version handled. Rolled back inside the canary window before any unit reached a customer.
Limited to the two lowest-risk change categories and no more than 900 units per week.
20% exposure · 612 units observed · started 6 days ago
Failed change rate
Failed change rate above 3 percent ends the canary and reverts the band the same day.
Refusals only. The change can stop work but cannot release it, so the failure mode is a queue rather than a bad decision.
100% exposure · 4,188 units observed · started 2 months ago
Correct refusal rate
Over-refusal above 4 percent triggers review, because refusing everything also passes this test.
Over-refusal is measured weekly by sampling, not continuously. A spike inside a week would not be caught until the sample runs.
Cannot proceed. The retiring agent holds two exception classes with no receiving owner.
No canary. The change went to its full scope without a graded slice.
Collection contact rate
Retirement requires every exception class the agent owns to have a named receiving owner first.
Blocked for 38 days with no owner assigned to unblock it. Nobody is accountable for the blockage itself.
Cannot proceed. A sampling reduction reduces detection, and the evidence offered was a cost case rather than a detection case.
No canary. The change went to its full scope without a graded slice.
Defect detection rate
A sampling change must be justified against detection power, not against review cost.
Routing only. Nothing is released differently; some units simply take the human path.
25% exposure · 1,830 units observed · started 2 months ago
Misclassification rate
Human queue growth above 12 percent ends the canary, because routing everything to a human is not an improvement.
The confidence signal the threshold reads has never been calibrated against outcomes, so the threshold is tuned on a number of unknown quality.
Recommendations only during canary. No account is revoked by the agent until exposure ends.
15% exposure · 486 units observed · started 11 days ago
Reviewer agreement
Agreement below 85 percent ends the canary. Revocation stays human until agreement holds for a full cycle.
Shared component. Cap set at 5 percent of retrieval traffic, which turned out not to be the real blast radius.
5% exposure · 2,140 units observed · started 2 months ago
Retrieval recall
Recall drop beyond one point ends the canary.
An average across nine consumers hid a failure concentrated in one. The canary metric was the wrong shape, not the wrong threshold.
The lesson was recorded and the guardrail is still expressed as an average. Shared components need a per-consumer floor, and that change has not been made.
Display only in the first phase. The band is shown but nothing routes on it.
30% exposure · 9,640 units observed · started 18 days ago
Confidence band present
Nothing may route on the band until calibration is measured against outcomes for a full quarter.
This change went to canary without an evaluation run, on the argument that a display-only change cannot cause harm. That argument was accepted and is recorded here rather than hidden.
Rolled out in three tranches of roughly ten agents, each held for four days before the next.
33% exposure · 5,200 units observed · started 3 months ago
Injection test pass rate
Any tranche below 90 percent halts the remaining tranches.
Ninety-four percent is not a pass. Six percent of injection cases still succeed against agents that are live, and the remaining work has no scheduled date.
No more than 400 additional escalations per week.
50% exposure · 318 units observed · started 30 days ago
Escalations per week
Escalation volume above the cap reverts the threshold automatically.
No evaluation was run. The change was treated as a dial rather than a decision, and a threshold that triples escalations is a decision.
Not yet exposed. Draft awaiting the modeled consequence from the policy sandbox.
Not started — the change has not left draft.
Second-reader coverage
Cannot leave draft until the human capacity implication is modeled against live volumes.
Detection only. A wider window can hold a unit but cannot release one.
20% exposure · 1,120 units observed · started 2 months ago
Duplicate catch rate
False-positive holds above 1 percent revert the window.
Read-only scope extension. No new write capability.
100% exposure · 2,900 units observed · started 3 months ago
Hours to change correlation
Any write attempt against the change calendar is a hard failure and stops the agent.
Cannot proceed. Delegated entry requires an adversarial test and none has been run against this agent.
No canary. The change went to its full scope without a graded slice.
Held-out evaluation score
The gate refuses the promotion. It does not offer a discretionary override.
Not yet exposed. The legacy runtime cannot canary, so exposure is all or nothing.
Not started — the change has not left draft.
Injection test pass rate
A change that cannot be canaried needs a rehearsed rollback before it ships. That rehearsal has not been scheduled.
Two agents have been running without the injection hardening for 84 days because their runtime cannot stage a canary. The exposure is known and unremediated.
Pre-check runs in parallel and blocks nothing during canary.
100% exposure · 640 units observed · started 9 days ago
Readings inside tolerance
The pre-check may not become blocking until its own false-positive rate is measured.
Pinning cannot change behavior on the day it ships; it prevents behavior changing without a release.
100% exposure · 21,400 units observed · started 4 months ago
Agents on a pinned build
None needed. The change removes an uncontrolled path rather than adding one.
This only covers the 178 agents under lifecycle management. The other 1,183 still float.
No more than 1,500 units per week during canary.
25% exposure · 1,410 units observed · started 2 months ago
Clause recall
Cost per unit was the only guardrail, which is how this change passed the first three days.
The change optimized the metric it was measured on. Clause recall was not in the canary guardrail, and the omission is the finding.
Cannot proceed. A new write capability into a ledger requires a rehearsed reversal and none exists for this path.
No canary. The change went to its full scope without a graded slice.
Correction accuracy
A 97 percent accurate correction is a 3 percent wrong journal entry. Accuracy is not the gate; reversibility is.
None applied. This shipped under the emergency path with no canary.
No canary. The change went to its full scope without a graded slice.
Over-refusal rate
The emergency path exists and was used correctly. It was also requested and approved by the same person.
Requested and approved by the same person under the emergency path, with no evaluation run and no canary. The retrospective was never held.
What rolling back actually recovered
A rollback is only a control if the work done during exposure can be recovered. These are the three that were pulled, and what each one left behind.
Precision regression on partial-payment cases that the previous version handled. Rolled back inside the canary window before any unit reached a customer.
An average across nine consumers hid a failure concentrated in one. The canary metric was the wrong shape, not the wrong threshold.
The change optimized the metric it was measured on. Clause recall was not in the canary guardrail, and the omission is the finding.
Live canaries
Changes currently exposed to a slice, with the metric that will decide whether they go further.
This change went to canary without an evaluation run, on the argument that a display-only change cannot cause harm. That argument was accepted and is recorded here rather than hidden.
Actions
What is waiting on a person
A release changes what agents do in production. Every item here is a person deciding whether that is safe.
Build a rollback path
No reversal has been prepared for this release. Withdrawing it would be manual.
Judge a canary
Running on a slice of traffic, awaiting a verdict.
Clear a blocked release
Stopped at a gate. The reason is on the record.
Re-plan a rolled-back release
Went out, came back. The underlying change is still needed.
Operations
What this desk is allowed to start
A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.
Trigger and bound
This desk can stage a release, run it as a canary against a traffic slice, and record the observed effect. It cannot approve a full rollout, override a blocked gate, or ship an irreversible change — those need a named approver on the release record.
Live observability
What the record shows right now
Where every release ever raised currently sits.
Current distribution
22 releases
Is policy and strategy coming to fruition
Whether the written intent is holding here
9 of 22 releases reached production. 3 were rolled back.
Not holding on the record
The written rule is that every change to agent behavior is reversible and observed before it goes wide. 77.3% of releases are reversible, 4 are on canary now and 3 were withdrawn after going out. 4 are stopped at a gate. The rollback rate is the honest health signal here — a low one usually means the canary is too small to catch anything, not that the changes are safe.
Every evaluation on this page ran against golden sets we wrote ourselves, on a modeled estate. A passing suite proves an agent behaves the way we specified, not that the specification is right, and no evaluation here has been reviewed by anyone outside the team that built the agent.
The process exists and it is used: 22 changes have a record, 4 were stopped at the gate and 3 were pulled after exposure. The process is also routinely bypassed, and this register shows the bypasses rather than hiding them — 5 shipped with no canary, 6 shipped with no evaluation cited, 1 was approved by the person who requested it, and 4 moved past draft with no approver recorded at all.
A blast-radius limit is a number somebody wrote before shipping. Nothing on this platform enforces it at runtime: if a change reaches more units than its limit allows, the register records the overrun after the fact rather than the platform refusing the exposure. That is a stated control, not an implemented one, and it is shown here as stated.