How the model got aligned

Simulation Revision Log

How the model got aligned — each round: the numbers, what they showed, why it happened, what we changed.

00

Reading the scores

A seam is the joint where one step hands work to the next — and where meaning can get lost in the transfer. The metric that matters is translation debt: how often meaning gets lost at those seams. Lower is better. A “win” is at least 5 points lower over the second half of the run.

01

Round 1 — baselinestraw man

Track B was a straw man. It had the phases and the committees but never touched the actual work. Track A got all the real mechanics.

Score (mock, 12 sprints): trust gap B−A = −0.73 — the test wanted B ahead and failed
What we saw

B “improved” nothing because it was wired to do nothing. The contest was decided before it started.

What we revised

Gave B real mechanics — it standardizes seams in waves, planning ahead, with success on each seam tied to how predictable that seam is, and backsliding under change fatigue. Made translation debt measured from the actual state of each seam — is the information actually there or not — instead of scanning the AI’s wording for keywords.

Why

A comparison with one working side isn’t a comparison. B needed to pull the same levers as A, just with different dynamics.

02

Round 2 — B has mechanicsmock

Score (mock, 16 sprints): translation debt A → 0%, B → 19%
What we saw

B now acts, but A still wins on the mock. We suspected the mock flattered A (its predictable “ask for the missing information” behavior feeds A’s evidence perfectly) and needed the real thing.

What we revised

Wired the sweep to real runs and added the knobs that decide the winner — how predictable the seams are, how good the vendor is, how long the horizon, how turbulent the environment, how slow the decision-making.

Why

The mock only exercises one scenario. The point was to see whether B could win under B-favoring conditions, and that needs the real model.

03

Round 3 — first real sweep6 runs

runA debtB debtwinner
Predictable · seed 194.622.6
Predictable · seed 345.723.4
Judgment · seed 263.225.5
Judgment · seed 113.126.6
Mixed · seed 131.921.5
Mixed · seed 303.621.9
What we saw

A won all six, even the “B should win” cell. B’s debt pinned at ~21–27% in every scenario, insensitive to every knob — the signature of a structural bias.

Mechanism (from the trace)

B’s waves standardized the wrong seams — it polished the high-risk-but-already-clean steps while missing the actually-broken predictable one (the coverage check) outright. And the seams it did fix backslid almost instantly (claim intake was fixed and reverted in the same sprint). A, meanwhile, fixed the actually-broken seams durably, one observed failure at a time.

What we revised (three bugs)
  1. B fixed the wrong seams — ranked by a static risk score, not by what was actually broken.
  2. B’s fixes washed off — backsliding on a flat timer, self-reinforcing.
  3. B got one chance — a missed seam stayed broken forever.
Why

Each fix is grounded in what real change management does: fix what’s broken, backslide under pressure rather than on a timer, keep working what slips.

04

Round 4 — fairness fixesre-running

runA debtB debtwinner
Predictable · seed 194.95.1tie
Predictable · seed 347.67.6tie
Judgment · seed 264.117.3
Judgment · seed 113.414.6
Mixed · seed 132.09.1
Mixed · seed 303.312.2
What we saw

B’s debt collapsed from ~23 to ~5 in the predictable cell — a tie, not an 18-point loss. B is now genuinely competitive where it should be, and A still wins the judgment cell. But the predictable cell ties rather than flips, and the mixed cell still goes to A on both runs.

Mechanism (from the trace)
  • Predictable: B completed all four broken seams (intake, coverage check, settlement, payment), retrying the ones it initially missed, with zero backsliding in the calm environment. That’s why it tied A.
  • Judgment: B hammered the seams it can’t see — the human negotiation step was missed 10 times — and the few it landed reverted. A’s evidence-built fixes held.
What this means (revised by Round 5)

The three fixes made B competitive — no longer a guaranteed loser — but didn’t reach a clean B-win. At the time we read this as an honest verdict: B is no longer rigged to lose, but it’s a contest B still doesn’t win. Round 5 showed that verdict was premature — the predictable cell tied instead of flipping because Track A still had a hidden head start we hadn’t found yet.

05

Round 5 — adversarial audit & de-riggingaudit

An independent, adversarial review audited the whole model with one question: is either side still favored by construction rather than by conditions? It found two places where the outcome was quietly written in — both pointing at Track A — plus a reproducibility bug.

What the audit found
  1. A could sharpen a seam it had never watched break. A’s rule is that it’s blind to failures it hasn’t observed — that blindness is its honest cost. But the code let A improve the worst AI seam every cycle whether or not it had seen that seam fail: a free, recurring head start that bypassed A’s own constraint. This is why the predictable cell tied instead of flipping to B.
  2. Employee trust-in-AI was hand-set to drift apart. Two tuned constants nudged B’s trust up and held A’s flat every sprint — the comment literally said they were “calibrated so B climbs and A holds.” The outcome was an input.
  3. Same seed didn’t reproduce. Three sources of randomness on A’s side were unseeded, so two identical runs could diverge.
What we changed

A can now only sharpen seams it has actually watched fail — kept the durability (its fixes still don’t wash off; that’s legitimate) and removed the omniscience. Trust now emerges for both sides from the same thing: what agents actually experience when the AI succeeds or fails. And every source of randomness is seeded.

Verified
  • Two same-seed runs now match within a single process (they didn’t before). Round 6 found this check too weak and hardened it.
  • Trust comes out close and earned — A 6.8, B 6.7 — instead of forced apart.
  • Both sides now share the same starting debt and separate only by conditions, not by construction.
Result — the real 6-run sweep (DeepSeek)
runA debtB debtΔ(A−B)winner
Predictable · seed 195.855.53+0.32tie
Predictable · seed 346.005.60+0.40tie
Judgment · seed 263.8815.29−11.41
Judgment · seed 112.232.23+0.00tie
Mixed · seed 132.208.37−6.17
Mixed · seed 304.0712.01−7.95
What actually happened — my prediction was wrong on both counts

I predicted the predictable cell would flip to B and the judgment cell would stay a clean A-win. Neither held. The predictable cell stayed a tie (though both seeds now lean B, where before they were even) — removing A’s crutch narrowed the gap but didn’t hand B the win. And the judgment cell went from a clean A-sweep to A + tie — one of A’s two tacit wins softened to a dead heat once its free head start was gone. Three A-wins, three ties, zero clean B-wins.

The honest read — a map, not a scoreboard

The point was never to make B win. It was to see whether the winner emerges from conditions. It does: A’s advantage concentrates where the work is tacit or mixed and turbulent (it wins the judgment and both mixed cells, by 6–11 points) and evaporates to a tie where the work is codifiable and calm. That A no longer wins everything — that a fake A-win collapsed to a tie the moment the crutch came out — is the de-rigging working. That B never gets a clean win is an honest limitation, not a rig: B’s a-priori design accuracy is capped on judgment seams it can’t template, so it converges to A on codifiable work and loses on tacit work. Disclosed, not hidden. The meaningful output is the regime map: ownership wins where meaning is tacit; the approaches converge where meaning is codifiable.

06

Round 6 — can someone else get the same numbers?reproducibility

A scientific result should be repeatable: run it again, get the same answer. A fresh review checked that — and found the sim was giving slightly different numbers each time it ran, even with everything set the same. Same settings, three runs, three answers — a wobble of about 12%:

same settings, 3 runsrun 1run 2run 3same answer?
before the fix129124139no
after the fix124124124
Why it happened

The sim decides “who leaves the company this quarter” by going down the list of employees and rolling the dice for each. The catch: the program was building that list in a slightly different order every time it started up. Same dice, dealt to people in a different order → different people left → every downstream number shifted. The randomness itself was fine; the order it got applied in wasn’t fixed.

What we changed

Put the employee list in a fixed, sorted order before rolling. Now the same dice always land on the same people, on any computer, every run.

Verified

We reproduced the wobble first, applied the one-line fix, then ran it several more times — identical numbers every time. All other checks still pass.

What this does and doesn’t change

This was only about repeatability — whether you and I get the exact same number from the same run. It never touched who wins A vs B or the “no longer rigged” verdict; within any single run both approaches always played by the same rules. It just means the published numbers can now be re-checked by anyone.

07

Round 7 — the trust number was lyingcorrectness

Score (2 canonical seeds, 16 sprints): translation debt A → 3.4, B → 17–19 · trust → ~4, not 8.0
What we saw

A full code review found a batch of correctness bugs quietly distorting the numbers. The loudest: trust was climbing to 8.0 in every run because an update formula multiplied and divided by the same count. Fix the arithmetic, and trust lands near baseline, ~4, where it belongs. The eleven profiles land on distinct values instead of converging.

What we revised

More than a dozen fixes. The exception rate had been counting “request info” and “override AI” as exceptions — it now counts only genuine escalations and lands near-equal across tracks. The discovery engine was handed an empty metric snapshot every sprint. Political friction only ever punished Track A. Underwriting and subrogation metrics lagged a sprint. THEARI advanced a phase early.

Attrition, and the fix to the fix

Attrition was ~5× too low, so nobody ever quit. Raising it made people leave — which emptied the claims department on one seed and stalled the pipeline to zero decisions. Departures now backfill a fresh hire, so churn stays visible without breaking the instrument.

The seed question, closed by math

The “two seeds is underpowered” objection got a power analysis instead of just more runs. The codifiable tie is parity — mean +0.19 with a confidence interval that straddles zero — not a hidden B-edge. The tacit win is decisive at −12. More seeds can’t flip either.

What this does and doesn’t change

The core finding survived and got sharper: B doesn’t escalate more, it leaks more. Translation debt is the decisive gap on both seeds. Trust, exceptions, and cost all re-baselined on two fresh canonical seeds.

08

What this teaches us about a better Track A

Collected across the rounds, the sim points at a specific design question:

  1. A’s edge is durability, not intelligence. A fixes the root cause — the seam where meaning is lost — and the fix sticks because the team built it from their own evidence. B fixes the symptom (a template, a dashboard) and it washes off. The winning ingredient is ownership, not speed.
  2. B’s edge is breadth, not accuracy. B covers everything fast but is blind to the judgment seams and can’t hold what it takes. A sees clearly but moves one seam at a time.
  3. The next model is the question between them: can you get A’s durability with B’s breadth? Fixes built from evidence that land fast and cover more than one seam per two sprints. The current sim is the instrument for testing that — the traces show exactly where A leaks time (slow learning, one-at-a-time throughput) and where B leaks trust (backsliding).
09

What this teaches us about a better Track B

The same trace reads the other way. B’s failures aren’t random; each one names a specific fix.

  1. Keep the breadth, aim it. B’s structural edge is covering everything fast. Its mistake was ranking by a static risk score and polishing already-clean seams while the broken ones leaked. A better B keeps the big-bang coverage but spends it on the seams where debt actually lives — the broken ones.
  2. Actually do Consolidate and Anchor. Kotter’s model has two whole phases for re-sustaining change that slips. B’s implementation skipped them — it fired a flat backslide timer instead of following up. A better B measures adoption and re-anchors when it dips, instead of letting it wash off.
  3. Know what you can’t standardize. B is blind to the judgment seams — the trace shows it missing the human negotiation step ten times, hammering judgment work it can’t template. A better B would spend its upfront budget on the predictable seams where its design is accurate, and not impose a process template on negotiation and liability where it isn’t. Honesty about the limit is the fix.
  4. Measure absorption, not activation. B’s steering committee reviews licenses and training completions — vanity metrics. Usage means people have access; absorption means the org can carry the change. A better B wires the committee to the outcome (translation debt) so governance closes the loop instead of spinning.
10

The symmetry

The symmetry is the point. A better A leans toward B’s breadth. A better B leans toward A’s ownership. They’re two ends of one axis — speed versus absorption — and the best transformation is the one that refuses to choose.

Each round of revision tightens the model. The winner is the least interesting output; the how and why is what lets us build a better Track A, a better Track B, and a more faithful sim.

Authoritative full run (post-fix, 16 sprints, seed 42): see the results dashboard →