Function cockpit
Research & Development
Agents run research operations. Scientists own hypotheses and conclusions.
Phase: Phase 2 — Supervised execution
Connect scientific programs to resources, evidence and decision gates: retrieval, scheduling, data integrity, simulation and reproducibility are agent work; hypothesis selection, safety and scientific conclusions remain human.
Trust score
0
Touchless
0%
Human review
0%
Override rate
0%
Capacity released
20-35% research-operating capacity (no presumption that success rates improve)
Accountable human
Chief Scientific Officer
CoS agent: R&D Chief of Staff
Autonomy, trust, and throughput at a glance
Health
Is this function holding
Four readings with a printed rule behind each, then the outcome measures the function is judged on. A reading without a rule is decoration.
Autonomy against target
62%
12 points short of the 74% target
Trust score
81
below the 85 floor: no promotion is eligible
Touchless rate
52%
48% of items still take a person somewhere in the run
Override rate
12%
people are reversing the agents often enough that the ceiling is doing real work
Research cycle time
-22%38days
Question framed to evidence-complete synthesis
Failed reproduction rate
-3.16.2%
Internal results that could not be re-derived
Evidence freshness
+694%
Citations refreshed within the policy window
Experiment backlog
-1441runs
Authorized experiments awaiting laboratory capacity
Patent deadlines at risk
-42filings
Disclosure windows inside 30 days
Operations
What the function is running
The roster and where it sits on the ladder, the split between what runs alone and what a person still signs, and a slice of the floor as it stands.
Where this roster sits on the autonomy ladder
A0 assisted through A4 autonomous. Moving an agent up a rung is a governance decision, not a config change.
Agents
10
Active now
9
Tasks / 24h
3,754
Mean success
92%
Work mix, as it stands
How the function's volume divides between the agents and the people.
Runs end to end without a person
52%
cleared inside the ceiling, no queue, no signature
A person reads it before it clears
38%
evidence posted, a named reviewer signs
A person reverses the agent
12%
the agent proposed, the human decided otherwise
Capacity released
20-35% research-operating capacity (no presumption that success rates improve)
What runs without a person
Committed to autonomy inside a defined ceiling.
- Scientific, technical and patent evidence retrieval
- Experiment scheduling and resource coordination
- Dataset integrity, lineage and permission checks
- Simulation under approved parameters and notebook organization
- Invention intake and portfolio milestone reporting
What stays with people
Judgment, accountability, and anything a regulator would ask a human about.
- Hypothesis selection and research direction
- Physical experiment authorization and laboratory safety
- Clinical, ethical and safety decisions
- Scientific conclusions and external publication
- Patent strategy, program funding and termination
On the floor right now
A slice of live work. The full board carries every item.
- RND-1140Autonomous
Evidence synthesis — polymer degradation pathways
Incubate · R&D CoS · 24m old · 412 sources
- RND-1141Awaiting human
Protocol authorization — temperature cycling study
Intensify · Safety officer · 3h 6m old · 18 runs
- RND-1142Escalated
Dataset license remediation — catalysis corpus
Incubate · Research governance · 1h 14m old · 3 studies
- RND-1143Autonomous
Parameter sweep — membrane permeability model
Intensify · R&D CoS · 52m old · 2,400 runs
- RND-1144Drafted
Invention record — composite bonding method
Industrialize · General Counsel · 7h 8m old · 1 filing
- RND-1145Complete
Portfolio gate pack — Q3 stage review
Immerse · Portfolio board · 10h 40m old · 11 programs
Actions
What this function is asking a person to do
Ordered by size. Each figure is a live count from this function's own work and governance records, not a target.
2
Items awaiting a person
queued against a named human, clock running
2
Governance decisions pending
an agent stopped at a gate and asked
0
Items older than 48 hours
on the floor long enough to be a problem
1
Agents below 80 confidence
running, but not at a level that supports a promotion
Brakes available right now
What a named human can pull today to stop this function, without waiting for an engineer.
- Unlicensed or improperly sourced data in a study
- Irreproducible output promoted as a result
- Fabricated or unresolvable citation
- Unsafe protocol proposed or scheduled
- Leakage of unpublished invention material
Live observability
What has actually been decided, and where each agent stops
A dashboard that shows only outcomes hides the decisions that produced them. This is the governance record as written, and the ceiling every agent is held to.
Governance record
Most recent first. Each entry names the actor and the call.
Unlicensed dataset detected in active study
ContainedData Steward found a third-party dataset without an enterprise license inside an active catalysis study. Analysis halted and the dependency graph published for research governance.
escalation · Data Steward · materiality high
Protocol proposal awaiting safety review
PendingExperiment Designer proposed a temperature-cycling protocol outside the previously validated range. Held for principal investigator and safety officer authorization.
approval · Safety officer · materiality high
Reproduction check passed on 14 results
Auto-executedReproducibility Agent re-derived 14 internal results from source data and code with matching outputs. Evidence written to the research ledger.
notify · Reproducibility Agent · materiality low
Competing patent filing overlaps active program
PendingInvention Agent detected claim overlap with a filing published this week. Disclosure timeline and novelty evidence packaged for Legal and the CSO.
escalation · General Counsel · materiality high
Escalation ceilings
Past the line the agent stops and hands the decision to the named human with the evidence attached.
- A3
R&D Chief of Staff
Ceiling: Coordination only — no scientific or funding authority
Then: Any program crossing a stage gate requires the named scientific owner before progression.
- A3
Evidence Scout
Ceiling: Licensed sources only; no paywall circumvention
Then: Any citation that cannot be resolved to a verifiable source is withheld and reported.
- A2
Hypothesis Mapper
Ceiling: Structuring only; never asserts a scientific conclusion
Then: Contradictory high-quality evidence is surfaced to the principal investigator unresolved.
- A2
Experiment Designer
Ceiling: Approved method library only; no novel protocol authorization
Then: Physical experiments, hazardous materials and protected subjects require named human authorization.
- A3
Lab Scheduler
Ceiling: Scheduling within authorized capacity only
Then: Resource conflicts affecting a critical-path program escalate to the program lead.
Is policy and strategy coming to fruition
Does the intent above this function reach the work inside it
Counted per sub-function, where each step only counts if the step before it did. A policy that never reaches a running workflow has not landed, however well it reads.
Chain from stance to transaction
4 of 11 sub-functions carry a ratified position, an active policy and a live workflow
Ratified position
4
of 11 sub-functions — a stance the leadership signed, not a draft
...and an active policy
4
a rule in force, with a version and an approver
...reaching a live workflow
4
the rule reaches something that actually runs
...and landing in a GBS tower
3
the run is executed on the shared transaction spine
7 of the 11 sub-functions are executing at volume without the full chain behind them. They run to a general standard rather than to a rule with a version and an approver, and the gap first appears at the position step.
What this page is, and what it is not
Every reading here is computed live from this function's own records, which makes it exact and makes it narrow. Volumes, unit costs, autonomy and trust are modeled: none has been reconciled against an enterprise resource planning system, a service management tool or a payroll register. Read the outcome measures as the shape of the argument, not as an audited result.
BCG cites a biopharma example with 35% R&D-efficiency improvement and 25% shorter cycles. Accenture estimates $180-240 billion in potential annual value from agentic twins across biopharma R&D and manufacturing, but this is a forward-looking sector estimate, not realized average value. A prudent benchmark is 20-35% improvement in research-operating capacity.