Plain English — a companion to “How It Works”

Is the Simulation Fair?

A pretend company can be built to say anything you want. Here’s how we keep this one honest — and how we know it’s grounded in the real world, not made up.

01

The trap every simulation falls into

Here’s the uncomfortable truth. If you build a computer model of a company and you’re quietly hoping one side wins, you will almost certainly get that side to win — without ever meaning to cheat. A number nudged here, an advantage granted there, and the model dutifully hands back the answer you were rooting for. Then you show it around as “proof.”

That’s not proof. That’s a mirror.

So the real question isn’t “did our preferred approach win?” It’s would this simulation have been willing to tell us we were wrong? A model that can only confirm your hopes proves nothing. A model that fights back, that produces results you didn’t want, is the only kind worth trusting. Everything below is how we force this one to be that second kind.

02

Six ways we keep it from cheatingPart one · Fair

“Fair” means neither team gets a secret advantage, and the model is genuinely allowed to reach either answer. Six safeguards enforce that.

Guard 1Identical
luck
Both teams get the identical luck. Risk If one team happened to get easier customers or luckier breaks, its win would mean nothing. Check Whether the AI succeeds or stumbles on any given handoff is decided by a fixed fingerprint of the exact claim, step, and moment — not a live coin flip. Both teams hit the identical success and failure on the identical work at the identical time. There is no luck to be unequal; any lasting difference can only come from their approach.
Guard 2Symmetric
scoreboard
The scoreboard is symmetric. Risk A scoring rule tilted even slightly toward one side quietly decides the winner before the race. Check The winner is whoever has clearly less cleanup work over the settled second half of the run. The rule is direction-blind: the same margin is required to declare either team the winner. It doesn’t know or care which team we like.
Guard 3No punching
bag
The losing team isn’t a punching bag. Risk The easiest way to rig a contest is to make the other side stupid — a cartoon of clueless executives staring at dashboards. Check The “big plan” team runs on the real playbooks real companies use (Kotter’s 8 Steps, Prosci ADKAR, McKinsey 7S), executed competently. A separate fidelity auditor inspects both teams every round, checking neither has drifted into caricature — and can halt the run if it finds too many violations. A weak opponent isn’t allowed to stand.
Guard 4Try to make
them win
We actively try to make the other side win. Risk If you never test whether the underdog can win, you never learn whether the deck was stacked. Check We deliberately build the conditions where the “big plan” team should win — the easy, well-documented, calm situations that favor planning over learning-by-doing — and check whether the model lets it. If it can produce that win, it isn’t rigged. If it never can, that’s a real finding worth chasing, not a design we hide.
Guard 5Hire a
skeptic
We hire a skeptic to attack it. Risk The person who built the model is the last person likely to spot their own thumb on the scale. Check We run adversarial audits whose only job is to find hidden advantages. One recently caught a real one: the “learn as you go” team was getting an unearned, permanent, never-fading discount on the exact number the whole score is built from. It was removed — and the entire result re-checked with it gone.
Guard 6Write the
prediction first
We write the prediction down first. Risk If you decide what counts as success after seeing the result, you’ll always find a way to call it a win. Check Before a re-run we commit the prediction in writing — “this specific situation should flip to this team.” The random starting conditions are chosen blind to the outcome, screened only for the weather we wanted, never for who wins. No moving the goalposts after the ball is in the air.
03

Why the pretend company behaves like a real onePart two · Realistic

Fair isn’t enough. A perfectly fair model of a company that doesn’t exist still tells you nothing. So the model is built on real numbers and real behavior, not invented ones.

70,000
Employees at the real company it’s modeled on — Progressive Insurance.
$86B
In real annual premiums, from public filings.
55%
Of claims the real company already runs through AI — the starting point.
87.3%
Combined ratio — the real profitability number we calibrate to.
3
Real insurance workflows: auto claims, commercial underwriting, subrogation.
43
Measured numbers across 6 dimensions, checked by 8 validity gates.

Real financials, not guesses. Size, money, and starting AI usage come from public filings, Glassdoor, and job postings. Real work, not a toy. Claims travel the actual chain of desks an insurer uses, across three genuine lines of business. Real people, not robots. Staff are modeled with personality, the Kahneman-and-Tversky biases, what drives them, and who they listen to — organizations run on people protecting turf and quietly fixing things with unwritten knowledge. Real playbooks, not straw men. Both approaches rest on published, field-tested methods, cross-checked against outside expert frameworks.

04

What actually counts as proof

Proof has nothing to do with our preferred approach winning. It’s whether the model captures enough reality that its output is evidence instead of decoration. Five plain tests:

It surprises
us
If every result just confirms what we already believed, we built an echo chamber. Real proof looks like discoveries we didn’t program — a bottleneck showing up somewhere unexpected. The model logs these on its own.
Practitioners
recognize it
If a claims executive reads a character’s decision log and says “I’ve managed three of that person,” the model caught something true.
It’s robust,
not fragile
Run it many times with different random starts. The big patterns hold; only the small details shift. If a tiny change flips the whole result, nothing was proved — which is exactly why we run across seeds instead of trusting one.
The losing side still
does something right
Even a clumsy transformation improves something. If the other team is a total punching bag, it’s propaganda, not evidence.
The counterfactual
is honest
We can point to conditions where the other approach would be the right call — and the model can actually produce that. A model that can only reach one answer isn’t measuring; it’s asserting.
05

The best evidence — it refused to give us the answer we wanted

Here’s the part most demos hide. When the skeptic’s audit removed that hidden advantage and we re-ran everything honestly, the model did not hand us a clean, satisfying story. It gave us an awkward, more believable one.

What the honest re-run actually showed. In the calm, easy situations that were supposed to favor careful up-front planning, the two approaches came out essentially tied — a hair’s difference, well inside the noise. In the messy, high-pressure situations, the “learn as you go” team won decisively. Neither result is the tidy landslide anyone was hoping for. That’s exactly what makes it trustworthy.

A rigged model gives you a clean win for your favorite. An honest model gives you a tie where the theory was weak and a clear result where it was strong — and makes you sit with the discomfort. We even ran the most extreme test we could: stacking every advantage in the underdog’s favor — perfect tools, the calmest conditions, the longest runway — to see if there was any honest way for it to win outright. There wasn’t. It came out a dead tie. That’s not the model being stubborn; that’s the model being honest.

The whole point in one line. We didn’t build this to win an argument. We built it to find out if we were wrong — and we wired it so it could tell us so. The day it stopped agreeing with us is the day it started being worth something.