LensReading which lens this session carries.
Research & Development/Experiment Design

Sub-function

Experiment Design

Owns protocol work end to end inside Research & Development. Runs 3 lines across 16 stations, drains the repeatable share into the C2L tower, and holds 4 human gates.

orchestrator Experiment Design Orchestratorhuman owner Director, Research Methods routes into c2l

Estate rating

BBB

71 / 100

autonomy 59%cap 72%

Protocol drafting and power analysis are agentic. Ethics clearance and scientific validity judgments are personally owned by a qualified researcher.

Agents

8

Lines

3

Units / wk

10,635

Timetable vs actual

22.3h → 19.4h

Cost / unit

$17

was $45

Run-rate saving

$15.7M

annualized at this volume

The 8 agents running this sub-function

Grouped by what they are for, not by where they sit. The agent org is process-shaped.

OrchestratorOwns the queue and arbitrates between desks.

Experiment Design Orchestrator

Owns the experiment design lane end to end. Sequences the other agents, holds the timetable, and decides what surfaces to a human.

ObserveA315 / 24h87%4% ovr$13/task

ceiling A3 · escalates to Escalates to Director, Research Methods when the unit falls outside the decision envelope or confidence drops below the floor.

AI GBSRuns the shared transaction spine.

Protocol Drafting Agent

Runs the repeatable portion of experiment design inside the shared transaction spine at spine cost and spine controls.

AI GBSA32,985 / 24h93%4% ovr$0.31/task

ceiling A3 · escalates to Escalates to Director, Research Methods when the unit falls outside the decision envelope or confidence drops below the floor.

TaskExecutes stations on a line.

Power Analysis Agent

Executes the experiment design step it owns, posts its evidence, and hands the unit to the next stage.

WorkflowA3291 / 24h89%4% ovr$0.44/task

ceiling A3 · escalates to Escalates to Director, Research Methods when the unit falls outside the decision envelope or confidence drops below the floor.

Control Design Agent

Executes the experiment design step it owns, posts its evidence, and hands the unit to the next stage.

WorkflowA34,039 / 24h98%3% ovr$0.16/task

ceiling A3 · escalates to Escalates to Director, Research Methods when the unit falls outside the decision envelope or confidence drops below the floor.

ObserveWatches signals and watches the agents.

Bias Detection Agent

Watches experiment design continuously. Detects drift, quantifies it, and routes what matters without waiting for a reporting cycle.

ObserveA393 / 24h88%16% ovr$11/task

ceiling A3 · escalates to Escalates to Director, Research Methods when the unit falls outside the decision envelope or confidence drops below the floor.

PolicyAuthors and versions the binding rules.

Protocol Compliance Agent

Holds the rules that bind every other agent in experiment design. Versions them, tests them against live decisions, and blocks work that breaches them.

PolicyA284 / 24h92%10% ovr$5.66/task

ceiling A3 · escalates to Escalates on any decision that would breach a live policy version, with the clause cited.

Design Policy Agent

Holds the rules that bind every other agent in experiment design. Versions them, tests them against live decisions, and blocks work that breaches them.

PolicyA29 / 24h94%13% ovr$11/task

ceiling A3 · escalates to Escalates on any decision that would breach a live policy version, with the clause cited.

ChallengerPaid to disagree before a human has to.

Experiment Design Challenger Agent

Adversarial reviewer for experiment design. Argues the opposite case on every position and flags where the evidence does not carry the claim.

StrategyA285 / 24h88%6% ovr$8.08/task

ceiling A3 · escalates to Escalates when a ratified position is still running against a break condition it flagged.

Estate rating breakdown

Seven control dimensions. Judgment readiness runs against autonomy on purpose.

Control Coverage74
Evidence Quality69
Override Discipline76
Data Integrity66
Recovery Readiness72
Cost Transparency71
Judgment Readiness70

People

What the humans stopped doing, and what they do now.

headcount on this work59.738.4 FTE

Redeployed into agent supervision, calibration and coaching.

inference cost per unit$2.29

Lines running in this sub-function

Each line is a workflow. Click a line to walk it station by station.

12 clear2 evidenced3 held
RD3AProtocol DevelopmentTransactional · 7 stations

Frame question

27 q

Draft protocol

8 q

Power analysis

3 live

Bias review

45 q

Scientist review

9 q

Ethics clearance

2 live

Register

17 q

RD3BProtocol DeviationTransactional · 5 stations

Detect

1 live

Classify

1 live

Assess validity impact

2 live

Scientist decide

1 live

Document

3 live

RD3CDesign Standard RefreshTransactional · 4 stations

Mine deviations

40 q

Draft standard

1 live

Council approve

2 live

Publish

1 live

Strategy positions this sub-function holds

An agent that only executes is a robot. These are the calls it made, the argument against each one, and what would prove it wrong.

RD3-POS-100

Experiment Design Operating Position

Ratifiedv4

Should protocol work stay in the shared transaction spine, or come back inside the function where the context lives?

Hold the ceiling at A3 for experiment design. Protocol drafting and power analysis are agentic. Ethics clearance and scientific validity judgments are personally owned by a qualified researcher. Raise it only after two consecutive quarters where the override rate stays under five percent and every override has a written cause.

confidence

64%

The challenge — Experiment Design Challenger Agent

Experiment Design Challenger Agent argues the recommendation leans on 3 quarters of data from a period with no volume shock. The confidence band overlaps the alternative, and the position does not say what it would take to be wrong. Recorded as a dissent, not a block.

Break conditions

  • ·Override rate rises above 16 percent for two consecutive months
  • ·Cost per protocol stops falling while volume keeps rising
  • ·A control failure in this lane reaches a customer or a regulator

Evidence gaps

  • ·No external comparator on a like-for-like unit definition
  • ·Exception cases under $6k are sampled, not fully measured
Alternative consideredCostRiskVerdict
Hold the current position$333k run rateKnown and priced. Cedes ground if comparators move faster.not selected
Raise the ceiling one level now$1477k to build controlsOverride rate is 7 percent; raising the ceiling before that settles imports the error into production.rejected on evidence
Move the exception tail to the spine$284k transitionLoses local context. Rework risk on the cases that are hardest to recover.selected
authored by Experiment Design Orchestratorratifier Director, Research Methods3 stated assumptions
outcome held

Ratified 2 months ago. Measured since: cost per protocol fell 38 percent against a forecast of 37 percent.

Policy register

The rules the agents above are bound by. Version, owner, approver, and the agents each rule constrains.

RefRuleScopeOwner agentApproved byVersionStatus
RD3-POL-200

Experiment Design Decision Envelope

Agents in this lane may act without a human when the protocol sits inside the stated value, risk and confidence envelope. Outside it, the unit holds at a gate with a named approver and a running clock.

binds 3 agents · EU · UK

envelopeProtocol Compliance AgentDirector, Research Methodsv1.6Active
RD3-POL-201

Experiment Design Evidence Standard

Every autonomous decision writes inputs, the rule version applied, the model and prompt version, the output and a reversal path. An action with no evidence record is treated as a control failure, not a fast decision.

binds 3 agents · EU · UK · US

evidenceProtocol Compliance AgentDirector, Research Methodsv1.3Active