Simulation Results
Run: full_20260820_122140_s42 · Started 2026-08-20T12:21:40 — Completed 2026-08-21 08:57
Executive summary
Track A — the organic, evidence-driven method — produced better outcomes on 7 of 8 metrics that showed meaningful divergence. Lower costs. Fewer errors. More customers stayed. But the real story isn’t the score. It’s the mechanism.
| What was measured | Track A | Track B | Winner | Plain English |
|---|---|---|---|---|
| Translation Debt Index | A 14.0 | B 23.4 | A | Track A loses less info at handoffs |
| Exception Rate | A 5.9 | B 6.5 | A | Track A hits fewer exceptions |
| Handoff Failure Rate | A 42.2 | B 62.5 | A | Track A has fewer broken handoffs |
| Supplement Request Rate | A 20.5 | B 22.8 | A | Track A requests fewer supplements — seams get fixed |
| Cost Per Claim | A 331.8 | B 376.6 | A | Track A spends less per claim |
| First Pass Accuracy | A 51.8 | B 36.2 | A | Track A gets more claims right the first time |
| Customer Retention After Claim | A 79.8 | B 76.6 | A | Track A keeps more customers after a claim |
| Policy Bind Rate | A 84.8 | B 86.2 | B | Track B closes more policies |
Track A delivered better business results across 7 of 8 metrics.
$332 vs $377 — 11.9% cheaper
Track A saves real money. Lower cost compounds across millions of claims.
51.8% vs 36.2% — 15.5-point gap · ~32% less rework
Rework is the hidden tax. Higher first-pass means your best people stop cleaning messes.
14.0 vs 23.4
The unpriced work of fragmented organizations — meetings, clarification, “quick syncs.” Why everything takes so long.
These three form a causal chain: lower translation debt → higher first-pass accuracy → lower cost per claim.
Bottom line: If you’re leading an AI transformation, stop asking “what’s our adoption target?” and start asking “what just broke, and what did we learn?” The operating model isn’t something you design in a conference room. It’s something you discover through implementation — one bounded deployment at a time.
The experiment
A controlled simulation compared two approaches to AI transformation inside a modeled U.S. property-and-casualty insurer over 12 quarterly sprints. Both tracks processed claims, underwriting, and subrogation with the same 13 psychologically differentiated agent profiles. The only difference was the transformation method.
What diverged
Metrics that showed meaningful differences between tracks (averaged across all 12 sprints).
| Metric | Track A | Track B | Δ | What it means |
|---|---|---|---|---|
| Translation Debt Index | 14.0 | 23.4 | −9.4 | Track A loses less information at handoffs |
| Exception Rate | 5.9 | 6.5 | −0.5 | Track B suppresses exceptions through governance filtering |
| Handoff Failure Rate | 42.2 | 62.5 | −20.2 | Track A has fewer handoffs that break |
| Supplement Request Rate | 20.5 | 22.8 | −2.2 | Track A requests fewer supplements — redesigns complete the seams |
| Cost Per Claim | 331.8 | 376.6 | −44.8 | Track A processes cheaper — less rework |
| First Pass Accuracy | 51.8 | 36.2 | +15.5 | Track A gets more claims right the first time |
| Customer Retention After Claim | 79.8 | 76.6 | +3.2 | Track A keeps more customers |
| Policy Bind Rate | 84.8 | 86.2 | −1.5 | Track B binds more policies |
What it means
The operating model for AI is not something an organization should fully design upfront — it is discovered through implementation.
- Track A’s observe-learn-redesign loop catches and fixes meaning-loss at organizational seams. Agents surfaced 137 missing-information gaps (Type 2 emergence) — decisions made from incomplete information where the incompleteness wasn’t recognized at decision time. Track A fixed these at the seam. Track B routed equivalent findings through a steering committee.
- Track A hits fewer exceptions — its loop fixes the seams that force claims into human handling. Track B’s higher exception rate means more claims stall on escalation and rework, not that Track B surfaces more problems.
- The trade-off is real: organic discovery produces better operational outcomes (lower costs, fewer errors, higher retention), but traditional change management provides more predictable governance (fewer surprises, higher trust scores, more policies bound). The right choice depends on whether you need speed of learning or predictability of process.
Why we know
Specific evidence from this 10,168-decision run.
- Agent behavior diverged. The same psychological profile produced different behavior across tracks because the organizational rules differed — e.g. Tommy used AI in 45% of Track A decisions vs 44% in Track B and overrode AI 13% vs 10%; Diana used AI in 4% vs 3%.
- Supplement requests: 20.5% vs 22.8% — a 2.2-point gap. Track A’s redesigns complete the seams, so its request rate declines as gaps are fixed; Track B never completes the seams, so its rate stays high.
- Exception rate: 5.9% vs 6.5%. Track A’s loop fixes the seams that force claims into human handling; Track B’s higher rate means more claims stall on escalation.
- Translation debt: meaning lost at handoffs in 14.0% of Track A decisions vs 23.4% in Track B. Track A addresses translation seams as they appear; Track B routes them through steering committee → roadmap → vendor assessment.
- THEARI discoveries: Track A has 137 discoveries, all validated by evidence across multiple sprints. Track B has 12 that auto-advanced on schedule — same count, zero evidence required.
- Sample size: across 10,168 total decisions (5,081 A, 5,087 B), the divergence patterns are consistent in direction across sprints — not artifacts of small samples.
Shareholder impact
Translated to a public-market lens, Track A’s approach is worth roughly $6.3B more in market capitalization than Track B — about $10.8 per share (5.1% on PGR’s $213 baseline). A directional estimate, not a forecast.
- Cost channel: Track A handles a claim for $258 vs $346 — $88 cheaper. At 5.5M claims/year that is $484M a year in savings, or 0.56 points of combined ratio.
- Top-line channel: Track A retains 4.5 more points of customers, compounding into ~1.4 points of premium growth and $184M more annual income.
- Combined: $668M/year more net income, $1.15 of EPS, and — at 9.4× — $10.8 per share ($6.3B market cap).
- Caveat: the simulation’s absolute cost scale is not calibrated to real dollars, so this is built from the A-vs-B gap (defensible), not absolute levels. Anchored to PGR’s actual Q2 2026 combined ratio of 87.3%.
Metrics — full divergence
All tracked metrics, A/B average and gap across the 12-sprint run. (Per-sprint time series for every metric live in metrics.csv, below; the headline translation-debt trajectory is charted on the Findings Report.)
| Metric | A | B | Δ | Favors |
|---|---|---|---|---|
| Translation debt | 14.0 | 23.4 | −9.4 | A |
| Exception rate | 5.9 | 6.5 | −0.5 | A |
| Handoff failure rate | 42.2 | 62.5 | −20.2 | A |
| Supplement requests | 20.5 | 22.8 | −2.2 | A |
| Cost per claim | 331.8 | 376.6 | −44.8 | A |
| Cycle time | 5.2 | 5.3 | −0.1 | tie |
| Trust polarization | 3.3 | 3.4 | −0.1 | tie |
| Underwriting cycle time | 4.2 | 4.2 | +0.0 | tie |
| Missed recovery opportunities | 12.6 | 12.6 | +0.0 | tie |
| First-pass accuracy | 51.8 | 36.2 | +15.5 | A |
| Customer retention | 79.8 | 76.6 | +3.2 | A |
| Risk selection accuracy | 97.8 | 98.0 | −0.2 | tie |
| Policy bind rate | 84.8 | 86.2 | −1.5 | B |
| Recovery dollars | 12,969 | 12,969 | +0.0 | tie |
| Employee AI trust | 4.2 | 4.1 | +0.1 | tie |
Discovery Engine outputs
LLM-driven pattern detection — finds both known patterns (18 categories) and novel patterns outside any predefined category. Novel patterns, 16 total: 13 social, 2 workflow, 1 unclassified.
Track A — discovered patterns (a register condensed by theme; every finding was evidence-validated across sprints):
ai_trust = 0.00 avoid AI despite positive experiences; boundary work hides AI output behind human judgment. Recommended: AI explainability, challenge mode, human-in-the-loop override authority.Track B — discovered patterns (timeline-gated, no evidence required):
fnol_intake, investigation_human, damage_estimation_human, coverage_verification_human, liability_determination_human, settlement_human, approval, payment, subrogation, escalated_to_kathryn — findings auto-advanced on schedule with zero evidence required.THEARI value-thesis validation
Tracks discoveries through five phases of increasing certainty. Track A is evidence-gated; Track B auto-advances on schedule.
Representative Track A findings (all high-confidence, evidence-validated): translation-debt spikes at the approval handoff; a severe trust bifurcation (high-trust, low-stress agents vs low-trust, high-exhaustion agents); a bottleneck migrating to a single human; cycle time reported as 0.0 (a metric defect the engine caught); repeated shadow-AI behavior among high-experience agents.
Capability state progression
Each track’s Discovery Engine unlocks higher autonomy as it proves reliability.
Agent behavior breakdown
| Agent | Decisions | AI use | Accept | Override | Escalate | Exceptions | Translation |
|---|---|---|---|---|---|---|---|
| AI Pipeline | 756 | 100% | 672 | 0 | 84 | 84 | 84 |
| Alicia | 743 | 74% | 438 | 29 | 198 | 198 | 0 |
| Tommy | 649 | 45% | 127 | 87 | 14 | 14 | 0 |
| Diana | 585 | 4% | 19 | 0 | 16 | 16 | 216 |
| Nick | 528 | 3% | 28 | 1 | 70 | 70 | 0 |
| Jordan | 400 | 60% | 238 | 0 | 28 | 28 | 0 |
| Greg | 345 | 79% | 254 | 17 | 4 | 4 | 0 |
| AI Underwriting | 278 | 100% | 269 | 0 | 9 | 9 | 9 |
| Rachel | 223 | 4% | 1 | 0 | 0 | 0 | 0 |
| Pat | 189 | 8% | 6 | 0 | 1 | 1 | 52 |
| Sanjay | 163 | 0% | 0 | 0 | 2 | 2 | 28 |
| AI Subrogation | 127 | 100% | 113 | 0 | 14 | 14 | 14 |
| Maria | 71 | 48% | 15 | 6 | 10 | 10 | 0 |
| Kathryn | 24 | 79% | 0 | 0 | 0 | 0 | 0 |
| Agent | Decisions | AI use | Accept | Override | Escalate | Exceptions | Translation |
|---|---|---|---|---|---|---|---|
| AI Pipeline | 754 | 100% | 656 | 0 | 98 | 98 | 98 |
| Alicia | 747 | 73% | 437 | 25 | 199 | 199 | 0 |
| Tommy | 668 | 44% | 125 | 65 | 10 | 10 | 0 |
| Diana | 569 | 3% | 11 | 0 | 23 | 23 | 404 |
| Nick | 526 | 4% | 23 | 2 | 77 | 77 | 0 |
| Jordan | 400 | 60% | 238 | 0 | 25 | 25 | 0 |
| Greg | 348 | 78% | 253 | 17 | 5 | 5 | 0 |
| AI Underwriting | 276 | 100% | 268 | 0 | 8 | 8 | 8 |
| Rachel | 220 | 3% | 1 | 0 | 0 | 0 | 0 |
| Pat | 183 | 8% | 8 | 0 | 1 | 1 | 53 |
| Sanjay | 163 | 0% | 0 | 0 | 3 | 3 | 56 |
| AI Subrogation | 123 | 100% | 109 | 0 | 14 | 14 | 14 |
| Maria | 70 | 49% | 17 | 4 | 9 | 9 | 0 |
| Kathryn | 40 | 80% | 0 | 0 | 0 | 0 | 0 |
Agent behavior divergence. Same agent profile, different behavior across tracks — the difference comes from track rules, not personality. The starkest: Diana’s translation incidents are 216 in Track A vs 404 in Track B (71% of her decisions).
Claim trace viewer
The same claims, processed through both tracks — where and why they diverge.
Handled by Pat, Diana, Kathryn, Maria across 10 steps. AI assisted at 5 decision points. Escalated 2 times. 2 exceptions (damage estimation ai, approval). Meaning lost at 5 handoffs.
Handled by the same four agents across 10 steps. AI assisted at 7 decision points. Escalated 5 times. 5 exceptions (payment, settlement human, damage estimation ai). Meaning lost at 5 handoffs.
Handled by Pat, Diana, Kathryn, Tommy across 8 steps. AI assisted at 3 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.
Handled across 10 steps. AI assisted at 5 points. Escalated 3 times; 3 exceptions; meaning lost at 7 handoffs.
Handled across 8 steps. AI assisted at 3 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.
Handled across 8 steps. AI assisted at 3 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.
Handled by Pat, Diana, Kathryn, Maria across 10 steps. AI assisted at 8 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.
Handled across 10 steps. AI assisted at 7 points. Escalated 2 times; 2 exceptions; meaning lost at 5 handoffs.
Handled by Diana, Maria, Tommy, AI Pipeline across 12 steps. AI assisted at 9 points. Escalated 2 times; 2 exceptions (liability determination ai); meaning lost at 5 handoffs.
Handled across 9 steps. AI assisted at 7 points; an agent overrode the AI once (Maria). Escalated 1 time; 1 exception; meaning lost at 3 handoffs.
Raw data
Download the underlying data files from this run.