The AI Transformation Experiment
Can the AI operating model be discovered through implementation, or must it be designed up front?
The question
of change
Organizations either overdesign the future state or underdesign the learning process. This experiment runs both theories head-to-head inside the same company.
A simulation (two identical copies of Progressive Insurance) spanning claims, underwriting, and subrogation.
The only way to settle the argument is to run it, not read another whitepaper.
The setup
Insurance
70,000 employees. $82 billion in premiums. 5 million auto claims a year. 55% of claims already flowing through AI. Massive public disclosure, a CEO who started as a claims adjuster, and a culture that shows up on Glassdoor.
A company model grounded in financials, press, job postings, and employee sentiment.
A simulation only matters if the starting point is plausible.
The tracks
not designed
Deploy AI into one bounded workflow without reorganizing. Watch what breaks. Extract the pattern. Redesign once. Repeat.
Compounding evidence instead of compounding assumptions.
The operating model emerges from contact with reality, not a steering-committee deck.
management
Steering committee. Vendor selection. Adoption dashboards. Resistance management. And now the mechanics to match Track A: a-priori seam standardization in waves, success scaled by how codifiable the seam is, reversion under change fatigue. The standard playbook, executed well.
A disciplined enterprise program with clear milestones — and real workflow mutations, not just narrative phases.
This is the comparison point. Track B pulls the same levers as Track A, with different dynamics.
Controlled chaos
market, same people
Both tracks face identical hurricanes, attrition waves, regulatory shifts, and market cycles. The agents don’t know which track they’re in.
A fair test of process, not luck.
If one track wins, it wins because of how it learns, not because the world was kinder.
Grounding
wears a source tag
10-K financials, Gainshare incentives, Glassdoor culture, OCEAN psychological profiles, Prosci / Kotter / McKinsey methodologies, Progressive’s disclosed AI pipeline, ACORD standards, and state DOI requirements. Assumptions get varied in sensitivity analysis.
A model where you can argue with the assumptions instead of hunting for them.
You can’t stress-test what you can’t see.
Architectureagents with psychology, inside a world engine
The architecture keeps one company model in memory, creates two identical copies, and lets each track’s process logic drive how change unfolds.
The sprint loopeleven steps, two tracks, one loop
Discovery is now a first-class citizen of every sprint. Each sprint runs the same sequence. The Discovery Engine reads the track’s sprint log and detects patterns with an LLM — no longer hardcoded threshold checks — while agents themselves surface missing handoff context from the incomplete information they observe.
What the experiment produces
Quarterly cycles of deployment, learning, and redesign across claims, underwriting, and subrogation.
Identical companies, different change processes.
Evidence-backed patterns extracted from the winning track.
Cycle time, loss ratio, and attrition divergence.