A controlled simulation, two tracks, one question

The AI Transformation Experiment

Can the AI operating model be discovered through implementation, or must it be designed up front?

View the results →

01

The question

QuestionTwo theories
of change
Most AI transformation fails before the first sprint.

Organizations either overdesign the future state or underdesign the learning process. This experiment runs both theories head-to-head inside the same company.

What you get

A simulation (two identical copies of Progressive Insurance) spanning claims, underwriting, and subrogation.

Why it matters

The only way to settle the argument is to run it, not read another whitepaper.

Measured cycle time, loss ratio, attrition, operating-model divergence
02

The setup

Why this companyProgressive
Insurance
The baseline is real.

70,000 employees. $82 billion in premiums. 5 million auto claims a year. 55% of claims already flowing through AI. Massive public disclosure, a CEO who started as a claims adjuster, and a culture that shows up on Glassdoor.

What you get

A company model grounded in financials, press, job postings, and employee sentiment.

Why it matters

A simulation only matters if the starting point is plausible.

Measured revenue, loss ratio, headcount, expense ratio
03

The tracks

Track ADiscovered,
not designed
Start small, learn fast, redesign from evidence.

Deploy AI into one bounded workflow without reorganizing. Watch what breaks. Extract the pattern. Redesign once. Repeat.

What you get

Compounding evidence instead of compounding assumptions.

Why it matters

The operating model emerges from contact with reality, not a steering-committee deck.

Measured discoveries per sprint, playbook growth, redesign cycle time
Track BTraditional change
management
Plan the future state, then roll it out.

Steering committee. Vendor selection. Adoption dashboards. Resistance management. And now the mechanics to match Track A: a-priori seam standardization in waves, success scaled by how codifiable the seam is, reversion under change fatigue. The standard playbook, executed well.

What you get

A disciplined enterprise program with clear milestones — and real workflow mutations, not just narrative phases.

Why it matters

This is the comparison point. Track B pulls the same levers as Track A, with different dynamics.

Measured adoption rate, milestone hit rate, ROI-estimate accuracy — plus translation debt, measured identically to Track A
04

Controlled chaos

Fair testSame storms, same
market, same people
The only variable is how change happens.

Both tracks face identical hurricanes, attrition waves, regulatory shifts, and market cycles. The agents don’t know which track they’re in.

What you get

A fair test of process, not luck.

Why it matters

If one track wins, it wins because of how it learns, not because the world was kinder.

Measured environmental-event impact, agent-state divergence
05

Grounding

EpistemicsEvery parameter
wears a source tag
Transparent epistemics, not hidden guesses.
PUBLIC FILINGPRESSINFERREDASSUMPTION

10-K financials, Gainshare incentives, Glassdoor culture, OCEAN psychological profiles, Prosci / Kotter / McKinsey methodologies, Progressive’s disclosed AI pipeline, ACORD standards, and state DOI requirements. Assumptions get varied in sensitivity analysis.

What you get

A model where you can argue with the assumptions instead of hunting for them.

Why it matters

You can’t stress-test what you can’t see.

Measured assumption count, sensitivity-sweep range
06

Architectureagents with psychology, inside a world engine

The architecture keeps one company model in memory, creates two identical copies, and lets each track’s process logic drive how change unfolds.

Orchestrator
Runs sprints, dispatches events, collects metrics, keeps both timelines in sync.
Track controllers
Track A: bounded workflow → deploy → observe → learn → redesign, driven by LLM discovery. Track B: committee → roadmap → rollout → adoption.
Discovery engine
LLM-driven pattern detection — finds friction and novel behaviors we didn’t pre-code. Agents surface missing handoff context; THEARI validates discoveries through five evidence-gated phases; capability unlocks as patterns prove out.
Agent runtime
LLM-powered roles with OCEAN profiles, biases, motivations, memory, and relationship graphs.
World engine
Market cycles, hurricanes, regulatory shifts, attrition — identical conditions for both tracks.
Company model
Departments, roles, workflows, incentives, metrics — the static ground truth.
Cross-cuttingTrace system
A single, self-contained audit trail captures every event across every layer above. One file, replayable start to finish. A Claude skill is available to query and analyze the trace output.
07

The sprint loopeleven steps, two tracks, one loop

Discovery is now a first-class citizen of every sprint. Each sprint runs the same sequence. The Discovery Engine reads the track’s sprint log and detects patterns with an LLM — no longer hardcoded threshold checks — while agents themselves surface missing handoff context from the incomplete information they observe.

1Events
Generate environmental events — identical for both tracks.
2Process
Execute track-specific process logic.
3Work
Process claims, underwriting applications, and subrogation cases through each track.
4Discover
Discovery Engine reviews the Track A sprint log — LLM-driven pattern detection across known categories and novel patterns.
5Gaps
Agent-surfaced gaps — agents surface missing handoff context; Track A completes the seam from observed-failure evidence, Track B standardizes it a-priori in rollout waves.
6Theari
Registers all discoveries; checks the five-phase gates (Theoretical → Empirical → Applicable → Replicable → Impact); records decision latency.
7Radiate
Outcome Radiator broadcasts Empirical+ discoveries to a cross-team feed — Track A unfiltered, Track B committee-gated.
8Unlock
Capability State checks the five-level unlock thresholds (Observe → Recommend → Bounded → Autonomous → Emergent).
9Update
Update dynamics (trust cascade, attrition contagion).
10Metrics
Collect metrics, including coordination tax and decision latency.
11Audit
Fidelity audit.
12Export
Export the trace, patterns, discoveries, and capability state.
08

What the experiment produces

Outputs
12 sprints

Quarterly cycles of deployment, learning, and redesign across claims, underwriting, and subrogation.

2 tracks

Identical companies, different change processes.

1 playbook

Evidence-backed patterns extracted from the winning track.

3 core metrics

Cycle time, loss ratio, and attrition divergence.