LensReading which lens this session carries.

Estate ratings and performance

One rating scale, four levels of zoom

Every agent rolls into a sub-function, every sub-function into a function, every function into the enterprise. The rating is not a satisfaction score. It is a judgment about whether this estate could be handed to an auditor tomorrow, and it is deliberately harsher than the autonomy number sitting next to it.

Health

How each function is performing

The measures every function scorecard draws on, pulled from the pages that own them.

Enterprise estate rating

A75/100

Sound, with a named weakness under active work.

Sub-functions rated154
Agents in scope1361
Tasks last 24h1,684,614
Mean success rate93%
Mean override rate7.3%
Agents trending down0
Open escalations46
Incidents on record16

The seven dimensions behind every rating

Enterprise mean across all rated sub-functions, with the weight each dimension carries

Control Coverage

75 · w 18%

Share of decision paths that pass through a named gate rather than around one.

Evidence Quality

75 · w 16%

Whether the trail an agent leaves would survive an external inspection unaided.

Override Discipline

83 · w 14%

How often humans reverse the agent, and whether the reversal is recorded with a reason.

Data Integrity

77 · w 14%

Confidence in the upstream records the agent is reasoning over.

Recovery Readiness

76 · w 14%

Time and cost to reverse a bad run, tested rather than assumed.

Cost Transparency

79 · w 10%

Whether cost to serve, including inference, is measured per unit rather than estimated.

Judgment Readiness

59 · w 14%

Capacity to handle the cases the rules do not cover. This score is deliberately lowest where autonomy is highest.

Rating distribution

How the 154 sub-function estates actually land

AA
2
A
69
BBB
78
BB
5

No estate is rated AAA. That is intentional. A perfect control rating on a two-year-old operating model would say more about the rater than the estate.

Function scorecards

Ranked by estate rating, not by autonomy. The two frequently disagree.

FunctionRatingScoreAutonomy now / capAgentsTasks 24hSuccessOverrideFallingOpen esc.IncidentsOpen the function record
A
77
78% / 88%90108,89993%7.5%031Line map
A
77
75% / 84%110141,91193%7.2%011Line map
A
76
66% / 77%90104,68293%7.6%07 ·11Line map
A
75
70% / 79%98112,47293%7.5%021Line map
A
75
74% / 83%88117,27193%7.1%021Line map
A
75
75% / 82%92111,74492%7%061Line map
A
75
74% / 81%91128,48493%7.1%02 ·11Line map
A
75
76% / 84%89110,06893%7.7%051Line map
A
75
66% / 76%111143,90493%7.1%04 ·11Line map
A
75
75% / 84%119134,87593%7.7%013Line map
BBB
74
72% / 78%99138,43692%7.1%021Line map
BBB
73
69% / 77%8692,61293%8.2%021Line map
BBB
73
61% / 69%106129,90193%6.8%021Line map
BBB
72
66% / 72%92109,35593%7.2%07 ·21Line map

The "open esc." column shows escalations still waiting on a person, with tier-1 count after the dot. A function can hold a strong rating and still carry escalations. The rating measures whether the estate is controlled, not whether it is quiet.

Ten estates to fix first

Lowest rated sub-functions across all 14 functions

Employment Legal

Legal · 8 agents · 10,567 units/wk

BB

51% / 56%

The lowest ceiling in the entire estate alongside Litigation. Employment law is jurisdiction-specific and personally consequential; agents prepare, qualified counsel always decides.

Network Design

Supply Chain · 8 agents · 7,324 units/wk

BB

56% / 60%

Deliberately capped low. Network modeling is fully agentic, but opening or closing a facility is a capital decision with employment consequences that only a board can take.

Discovery Research

Research & Development · 7 agents · 2,471 units/wk

BB

71% / 78%

Literature synthesis and hypothesis generation are exactly where agents outperform — they read everything. A scientist validates before anything becomes a program.

Innovation Portfolio

Research & Development · 8 agents · 11,992 units/wk

BB

58% / 64%

Agents assemble evidence and even recommend killing a program. Stopping or funding research is a capital decision with careers attached, so a committee owns it.

Environment Health & Safety

Administration · 11 agents · 8,482 units/wk

BB

50% / 54%

Deliberately the lowest ceiling in Administration. Where a decision can injure someone, a competent human signs. Agents accelerate investigation, never the verdict.

Brand & Positioning

Marketing · 9 agents · 7,593 units/wk

BBB

58% / 64%

Agents author the positioning with evidence and dissent attached. What the company says about itself is ratified by a named executive.

Territory & Quota Planning

Sales · 9 agents · 135 units/wk

BBB

46% / 66%

Segmentation and capacity math are fully agentic. Quota is a compensation commitment to named people, so the final allocation carries a human signature.

Account Planning

Sales · 9 agents · 4,088 units/wk

BBB

64% / 72%

Agents author the account strategy end to end. The rep owns the relationship and ratifies the plan they will actually run.

Commercial Advisory

Legal · 8 agents · 14,214 units/wk

BBB

63% / 64%

Agents draft the advice and the reasoning. Legal advice is privileged and relied upon, so a qualified human signs every answer that leaves the function.

Corporate & Governance

Legal · 8 agents · 7,927 units/wk

BBB

68% / 72%

Assembly, calendaring and filing mechanics are automated. Minutes and authority changes are approved by named officers because they are the legal record of what the company decided.

Actions

What is waiting on a person

A scorecard is only worth reading if somebody is expected to act on a red. These are this week’s reds.

Operations

What this desk is allowed to start

A surface that only reports is not operable. This is the work this page can set in motion, and the bound it runs into.

Trigger and bound

This desk can compose a scorecard: pull each measure from its owning page, apply the declared threshold and publish the result. It cannot change a threshold, override a rating, or exclude a measure — a scorecard here is a view over other pages, never a separate set of numbers.

Live observability

What the record shows right now

What is currently red across the measures every function scorecard draws on.

Current distribution

61 reds

Breached service levels1220%
Unresolved escalations4675%
Breached budget caps35%

Is policy and strategy coming to fruition

Whether the written intent is holding here

14 functions are scored against 61 ratified positions and 395 active policy items.

Nothing on the record settles this

The intent is a single scorecard per function that a leader can read in a minute and act on. Structurally that holds: every measure shown is drawn from the page that owns it, so there is no second version of a number anywhere on this platform. What it cannot yet do is weight: the scorecard treats a breached service level and an unratified strategy position as comparable signals, because no agreed weighting exists. Reading it as a ranking of functions would be wrong.

How to read this without fooling yourself

A high success rate on a high-volume transactional agent is close to meaningless on its own. The work is repeatable, so it should be near perfect. The number that matters there is override rate and the reason attached to each override.

Judgment Readiness is scored lowest exactly where autonomy is highest. That is the honest tension in this model: the more of the routine an estate has automated, the less practiced its people are on the cases that fall outside the rules.

Nothing on this page is a benchmark. These are this estate's own measurements against its own control set. Comparing a function to an external figure would require the same definitions on both sides, and those definitions rarely survive contact.