Results dashboard

Simulation Results

Run: full_20260820_122140_s42  ·  Started 2026-08-20T12:21:40 — Completed 2026-08-21 08:57

01

Executive summary

One approach discovered what worked. The other announced what would work.

Track A — the organic, evidence-driven method — produced better outcomes on 7 of 8 metrics that showed meaningful divergence. Lower costs. Fewer errors. More customers stayed. But the real story isn’t the score. It’s the mechanism.

What was measuredTrack ATrack BWinnerPlain English
Translation Debt IndexA 14.0B 23.4Track A loses less info at handoffs
Exception RateA 5.9B 6.5Track A hits fewer exceptions
Handoff Failure RateA 42.2B 62.5Track A has fewer broken handoffs
Supplement Request RateA 20.5B 22.8Track A requests fewer supplements — seams get fixed
Cost Per ClaimA 331.8B 376.6Track A spends less per claim
First Pass AccuracyA 51.8B 36.2Track A gets more claims right the first time
Customer Retention After ClaimA 79.8B 76.6Track A keeps more customers after a claim
Policy Bind RateA 84.8B 86.2BTrack B closes more policies

Track A delivered better business results across 7 of 8 metrics.

The three metrics that matter
Cost per claim

$332 vs $377 — 11.9% cheaper

Track A saves real money. Lower cost compounds across millions of claims.

First-pass accuracy

51.8% vs 36.2% — 15.5-point gap · ~32% less rework

Rework is the hidden tax. Higher first-pass means your best people stop cleaning messes.

Translation debt

14.0 vs 23.4

The unpriced work of fragmented organizations — meetings, clarification, “quick syncs.” Why everything takes so long.

These three form a causal chain: lower translation debt → higher first-pass accuracy → lower cost per claim.

The real differenceTrack A found real friction patterns and validated each with evidence before acting. Track B’s discoveries auto-advanced on schedule — same count of findings, zero evidence required. One track let reality correct the plan; the other defended the plan from reality. The difference showed up in outcomes, not sentiment: trust stayed roughly flat in both tracks, but Track A fixed its seams and cut handoff failures, exceptions, and cost per claim.

Bottom line: If you’re leading an AI transformation, stop asking “what’s our adoption target?” and start asking “what just broke, and what did we learn?” The operating model isn’t something you design in a conference room. It’s something you discover through implementation — one bounded deployment at a time.

02

The experiment

A controlled simulation compared two approaches to AI transformation inside a modeled U.S. property-and-casualty insurer over 12 quarterly sprints. Both tracks processed claims, underwriting, and subrogation with the same 13 psychologically differentiated agent profiles. The only difference was the transformation method.

Track A — Organic DiscoverySWT
Deploy AI into existing structures. Observe what breaks. Extract learning. Redesign one thing at a time. No reorg, no new roles, no RACI redesign.
Track B — Traditional Change MgmtKotter / Prosci
Build coalition. Form vision. Vendor selection. Phased rollout. Resistance management. Steering-committee governance.
Scale
10,168 total agent decisions across claims, underwriting, and subrogation domains.
03

What diverged

Metrics that showed meaningful differences between tracks (averaged across all 12 sprints).

MetricTrack ATrack BΔWhat it means
Translation Debt Index14.023.4−9.4Track A loses less information at handoffs
Exception Rate5.96.5−0.5Track B suppresses exceptions through governance filtering
Handoff Failure Rate42.262.5−20.2Track A has fewer handoffs that break
Supplement Request Rate20.522.8−2.2Track A requests fewer supplements — redesigns complete the seams
Cost Per Claim331.8376.6−44.8Track A processes cheaper — less rework
First Pass Accuracy51.836.2+15.5Track A gets more claims right the first time
Customer Retention After Claim79.876.6+3.2Track A keeps more customers
Policy Bind Rate84.886.2−1.5Track B binds more policies
04

What it means

The operating model for AI is not something an organization should fully design upfront — it is discovered through implementation.

  • Track A’s observe-learn-redesign loop catches and fixes meaning-loss at organizational seams. Agents surfaced 137 missing-information gaps (Type 2 emergence) — decisions made from incomplete information where the incompleteness wasn’t recognized at decision time. Track A fixed these at the seam. Track B routed equivalent findings through a steering committee.
  • Track A hits fewer exceptions — its loop fixes the seams that force claims into human handling. Track B’s higher exception rate means more claims stall on escalation and rework, not that Track B surfaces more problems.
  • The trade-off is real: organic discovery produces better operational outcomes (lower costs, fewer errors, higher retention), but traditional change management provides more predictable governance (fewer surprises, higher trust scores, more policies bound). The right choice depends on whether you need speed of learning or predictability of process.
05

Why we know

Specific evidence from this 10,168-decision run.

  • Agent behavior diverged. The same psychological profile produced different behavior across tracks because the organizational rules differed — e.g. Tommy used AI in 45% of Track A decisions vs 44% in Track B and overrode AI 13% vs 10%; Diana used AI in 4% vs 3%.
  • Supplement requests: 20.5% vs 22.8% — a 2.2-point gap. Track A’s redesigns complete the seams, so its request rate declines as gaps are fixed; Track B never completes the seams, so its rate stays high.
  • Exception rate: 5.9% vs 6.5%. Track A’s loop fixes the seams that force claims into human handling; Track B’s higher rate means more claims stall on escalation.
  • Translation debt: meaning lost at handoffs in 14.0% of Track A decisions vs 23.4% in Track B. Track A addresses translation seams as they appear; Track B routes them through steering committee → roadmap → vendor assessment.
  • THEARI discoveries: Track A has 137 discoveries, all validated by evidence across multiple sprints. Track B has 12 that auto-advanced on schedule — same count, zero evidence required.
  • Sample size: across 10,168 total decisions (5,081 A, 5,087 B), the divergence patterns are consistent in direction across sprints — not artifacts of small samples.
06

Shareholder impact

Translated to a public-market lens, Track A’s approach is worth roughly $6.3B more in market capitalization than Track B — about $10.8 per share (5.1% on PGR’s $213 baseline). A directional estimate, not a forecast.

  • Cost channel: Track A handles a claim for $258 vs $346 — $88 cheaper. At 5.5M claims/year that is $484M a year in savings, or 0.56 points of combined ratio.
  • Top-line channel: Track A retains 4.5 more points of customers, compounding into ~1.4 points of premium growth and $184M more annual income.
  • Combined: $668M/year more net income, $1.15 of EPS, and — at 9.4× — $10.8 per share ($6.3B market cap).
  • Caveat: the simulation’s absolute cost scale is not calibrated to real dollars, so this is built from the A-vs-B gap (defensible), not absolute levels. Anchored to PGR’s actual Q2 2026 combined ratio of 87.3%.
07

Metrics — full divergence

5,081
Track A decisions
5,087
Track B decisions
16
Novel patterns outside 18 known categories
137 / 12
THEARI discoveries A evidence-gated · B timeline-gated

All tracked metrics, A/B average and gap across the 12-sprint run. (Per-sprint time series for every metric live in metrics.csv, below; the headline translation-debt trajectory is charted on the Findings Report.)

MetricABΔFavors
Translation debt14.023.4−9.4
Exception rate5.96.5−0.5
Handoff failure rate42.262.5−20.2
Supplement requests20.522.8−2.2
Cost per claim331.8376.6−44.8
Cycle time5.25.3−0.1tie
Trust polarization3.33.4−0.1tie
Underwriting cycle time4.24.2+0.0tie
Missed recovery opportunities12.612.6+0.0tie
First-pass accuracy51.836.2+15.5
Customer retention79.876.6+3.2
Risk selection accuracy97.898.0−0.2tie
Policy bind rate84.886.2−1.5B
Recovery dollars12,96912,969+0.0tie
Employee AI trust4.24.1+0.1tie
08

Discovery Engine outputs

LLM-driven pattern detection — finds both known patterns (18 categories) and novel patterns outside any predefined category. Novel patterns, 16 total: 13 social, 2 workflow, 1 unclassified.

Track A — discovered patterns (a register condensed by theme; every finding was evidence-validated across sprints):

Adoption resistance / shadow AIhigh confidence · Diana, Pat, Sanjay, Tommy, Nick
Agents with ai_trust = 0.00 avoid AI despite positive experiences; boundary work hides AI output behind human judgment. Recommended: AI explainability, challenge mode, human-in-the-loop override authority.
Trust paradox — positive & negative experienceNOVEL · 3+ sprints persistence
Agents with the most positive AI experience land at either zero trust (Pat, Nick) or very high trust (Jordan, Greg); Tommy’s many positive runs still yield zero trust. Recommended: calibration training on realistic AI failure cases.
Boundary work — manual verificationmoderate-high confidence
Humans re-check and justify even accepted AI recommendations. Recommended: tiered verification (auto / spot-check / full review).
Bottleneck migration — to a single human95% conf · sprint 15
Work piles up on Diana (553 decisions, exhaustion 8.0) while Ron, Tricia, Mike, Leslie are idle. Recommended: workload balancing across available agents.
Metric inversion — cost per claimmoderate-high confidence
Cost stays high despite high AI auto-processing. Recommended: cost-benefit analysis of AI auto-processing vs human routing.
Exhaustion–trust disconnectNOVEL · 65% conf
Exhaustion and AI trust decouple — Diana exhausted with zero trust, Mike exhausted with high trust. Recommended: individual interviews and personalized interventions.
Decision-volume–trust paradoxNOVEL · 60% conf
No clean relationship between decision volume and trust. Recommended: audit AI routing so all agents are offered AI support.
Trust cascade — positive85% conf · 3 sprints
High-exposure agents (Jordan, Greg, Alicia) carry trust outward. Recommended: use them as AI champions for resistant agents.

Track B — discovered patterns (timeline-gated, no evidence required):

Agent-surfaced gaps — per seamReplicable / Impact · 12 findings
Missing handoff context surfaced at fnol_intake, investigation_human, damage_estimation_human, coverage_verification_human, liability_determination_human, settlement_human, approval, payment, subrogation, escalated_to_kathryn — findings auto-advanced on schedule with zero evidence required.
09

THEARI value-thesis validation

Tracks discoveries through five phases of increasing certainty. Track A is evidence-gated; Track B auto-advances on schedule.

Track A — Evidence-Gated
137 friction discoveries, 0 absence discoveries
Theoretical 49Empirical 2Applicable 16Replicable 10Impact 60
Track B — Timeline-Gated
12 friction discoveries, 0 absence discoveries
Theoretical 0Empirical 0Applicable 2Replicable 8Impact 2

Representative Track A findings (all high-confidence, evidence-validated): translation-debt spikes at the approval handoff; a severe trust bifurcation (high-trust, low-stress agents vs low-trust, high-exhaustion agents); a bottleneck migrating to a single human; cycle time reported as 0.0 (a metric defect the engine caught); repeated shadow-AI behavior among high-experience agents.

10

Capability state progression

Each track’s Discovery Engine unlocks higher autonomy as it proves reliability.

Track A — 140 detected, 5 validated, 5 successful redesigns, 0 incidents
ObserveRecommendBoundedAutonomousEmergent
Track B — 0 detected, 0 validated, 0 successful redesigns, 0 incidents
ObserveRecommendBoundedAutonomousEmergent
11

Agent behavior breakdown

Track A — Organic Discovery
AgentDecisionsAI useAcceptOverrideEscalateExceptionsTranslation
AI Pipeline756100%6720848484
Alicia74374%438291981980
Tommy64945%1278714140
Diana5854%1901616216
Nick5283%28170700
Jordan40060%238028280
Greg34579%25417440
AI Underwriting278100%2690999
Rachel2234%10000
Pat1898%601152
Sanjay1630%002228
AI Subrogation127100%1130141414
Maria7148%15610100
Kathryn2479%00000
Track B — Traditional Change Mgmt
AgentDecisionsAI useAcceptOverrideEscalateExceptionsTranslation
AI Pipeline754100%6560989898
Alicia74773%437251991990
Tommy66844%1256510100
Diana5693%1102323404
Nick5264%23277770
Jordan40060%238025250
Greg34878%25317550
AI Underwriting276100%2680888
Rachel2203%10000
Pat1838%801153
Sanjay1630%003356
AI Subrogation123100%1090141414
Maria7049%174990
Kathryn4080%00000

Agent behavior divergence. Same agent profile, different behavior across tracks — the difference comes from track rules, not personality. The starkest: Diana’s translation incidents are 216 in Track A vs 404 in Track B (71% of her decisions).

12

Claim trace viewer

The same claims, processed through both tracks — where and why they diverge.

Claim CL-02-0024
Track A — Organic Discovery

Handled by Pat, Diana, Kathryn, Maria across 10 steps. AI assisted at 5 decision points. Escalated 2 times. 2 exceptions (damage estimation ai, approval). Meaning lost at 5 handoffs.

Track B — Traditional Change Mgmt

Handled by the same four agents across 10 steps. AI assisted at 7 decision points. Escalated 5 times. 5 exceptions (payment, settlement human, damage estimation ai). Meaning lost at 5 handoffs.

How they differ: Track A surfaced fewer issues (2 exceptions vs 5; 2 escalations vs 5). Routine claims follow similar paths; divergence emerges on edge cases.
Claim CL-03-0018
Track A — Organic Discovery

Handled by Pat, Diana, Kathryn, Tommy across 8 steps. AI assisted at 3 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.

Track B — Traditional Change Mgmt

Handled across 10 steps. AI assisted at 5 points. Escalated 3 times; 3 exceptions; meaning lost at 7 handoffs.

How they differ: Track A had fewer exceptions (2 vs 3), fewer escalations (2 vs 3), and 5 translation incidents vs 7. What it means: Track A’s lower translation debt is the direct result of the observe-learn-redesign loop — meaning-loss at seams gets fixed, not documented and deferred.
Claim CL-05-0011
Track A — Organic Discovery

Handled across 8 steps. AI assisted at 3 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.

Track B — Traditional Change Mgmt

Handled across 8 steps. AI assisted at 3 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.

How they differ: Both tracks handled this claim similarly — routine claims don’t expose the rule differences. What it means: divergence emerges on edge cases, not the routine flow.
Claim CL-02-0006
Track A — Organic Discovery

Handled by Pat, Diana, Kathryn, Maria across 10 steps. AI assisted at 8 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.

Track B — Traditional Change Mgmt

Handled across 10 steps. AI assisted at 7 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.

How they differ: Both tracks handled this claim similarly.
Claim CL-01-0000
Track A — Organic Discovery

Handled by Diana, Maria, Tommy, AI Pipeline across 12 steps. AI assisted at 9 points. Escalated 2 times; 2 exceptions (liability determination ai); meaning lost at 5 handoffs.

Track B — Traditional Change Mgmt

Handled across 9 steps. AI assisted at 7 points; an agent overrode the AI once (Maria). Escalated 1 time; 1 exception; meaning lost at 3 handoffs.

How they differ: Track B agents override more — possibly a narrower pre-vetted AI scope or governance requiring human sign-off on more steps. Track A’s higher exception count can be a signal of organizational health — issues are visible; Track B’s lower count may mean issues exist but aren’t flagged. What it means: Track B’s lower translation debt here may reflect a pre-designed RACI preventing ambiguity — at the cost of flexibility.
Why they differ — the rules
Interruption design
Track A: pause, override, and stop are designed in; agents exercise them without political consequence. Result: more exceptions surfaced, more supplements requested. Track B: exceptions are escalations that signal failure; governance filters before visibility. Result: fewer surfaced, but handoff failures accumulate.
Learning extraction
Track A: after every deployment the Discovery Engine finds patterns and proposes one redesign; translation debt is addressed as it appears. Track B: learning goes through steering committee → roadmap phase → vendor assessment; debt accumulates until the “consolidate gains” phase.
Deployment constraint
Track A: deploy into existing structures; the operating model is discovered through implementation. Track B: reorg first (McKinsey 7S), train first (ADKAR), build coalition first (Kotter Step 2); the model is designed before implementation.
THEARI gating
Track A: evidence-gated — a discovery advances only with enough evidence sprints, confidence, and measured improvement. Track B: timeline-gated — discoveries auto-advance after N sprints regardless of evidence.
13

Raw data

Download the underlying data files from this run.

Metrics CSV
Time-series metrics for both tracks — 22 metrics across all sprints. Download metrics.csv
Discovery patterns
All patterns detected by the Discovery Engine, both tracks. Download discovery_patterns.json
THEARI register
Phase progression for every discovery — evidence-gated (A) and timeline-gated (B). Download theari_register.json
Novelty register
Emergent patterns outside predefined categories — what we didn’t predict. Download novelty_register.json
Capability state
5-level unlock progression for each track’s Discovery Engine. Download capability_state.json
Decisions JSONL
Every agent decision — claim ID, step, reasoning, AI usage, exceptions. Download decisions.jsonl
Full state
Complete simulation state — agent summaries, trust, playbook, governance. Download state.json
Audit trace
Comprehensive event log — every decision, discovery, absence, trust shift, and environmental event in one file. Download trace.jsonl