Robust, not fragile
Cross-Run Results
Every simulation run side by side — which outcome patterns hold across seeds and which are per-run noise.
01
Executive summary
12 of 14 outcome metrics hold across every run — 10 for Track A, 0 for Track B, 2 tied.
Track A owns Translation debt, Exception rate, Handoff failure rate, Supplement requests, Cost per claim, Cycle time, Trust polarization, First-pass accuracy, Customer retention, and Policy bind rate. Missed recovery opportunities and Recovery dollars are a dead heat. Underwriting cycle time and Risk selection accuracy flip between runs — real noise, not signal.
02
Metric winners by runseed 42 · 43 · 44
Three runs compared. A pattern is marked stable when the same track wins it in every run. If it holds across seeds, the result is structural, not a fluke.
| Metric | Seed 42 | Seed 43 | Seed 44 | Stable |
|---|---|---|---|---|
| Translation debtMeaning lost when work passes between steps.Track A loses less information at handoffs | 25 vs 28A | 25 vs 31A | 29 vs 32A | ✓ |
| Exception rateHow often work hits an exception needing a human decision.Track A hits fewer exceptions | 15 vs 25A | 19 vs 30A | 23 vs 29A | ✓ |
| Handoff failure rateHow often a handoff between steps breaks.Track A has fewer broken handoffs | 40 vs 50A | 46 vs 58A | 49 vs 53A | ✓ |
| Supplement requestsHow often agents ask for missing information.Track A requests fewer supplements | 64 vs 98A | 52 vs 112A | 81 vs 132A | ✓ |
| Cost per claimDollars to process a claim, including rework.Track A spends less per claim | 400 vs 434A | 385 vs 434A | 501 vs 542A | ✓ |
| Cycle timeDays from claim start to finish.Track A finishes claims faster | 5 vs 5A | 4 vs 5A | 7 vs 7A | ✓ |
| Trust polarizationHow divided employees are about the AI.Track A has less internal division about AI | 3 vs 3A | 2 vs 3A | 3 vs 3A | ✓ |
| Underwriting cycle timeTime to issue a policy.Track A issues policies faster | 4 vs 4B | 4 vs 4B | 4 vs 4A | — |
| Missed recovery opportunitiesRecovery chances subrogation left on the table.Track A misses fewer recovery chances | 8 vs 8= | 5 vs 5= | 3 vs 3= | ✓ |
| First-pass accuracyShare of claims handled right the first time, no rework.Track A gets more claims right the first time | 57 vs 45A | 52 vs 37A | 28 vs 21A | ✓ |
| Customer retentionShare of customers who stay after a claim.Track A keeps more customers | 80 vs 76A | 79 vs 74A | 73 vs 71A | ✓ |
| Risk selection accuracyHow accurately risks are priced.Track A prices risks more accurately | 82 vs 92B | 89 vs 87A | 87 vs 88B | — |
| Policy bind rateHow often a quoted policy is bound.Track A closes more policies | 90 vs 85A | 88 vs 87A | 85 vs 84A | ✓ |
| Recovery dollarsDollars recovered through subrogation.Track A recovers more dollars | 4199 vs 4199= | 4921 vs 4921= | 5394 vs 5394= | ✓ |
| Employee AI trustHow much employees trust the AI (0–10).Trust stays roughly flat — the divergence is in outcomes, not sentiment | 4.9 vs 5.3B | 5.0 vs 5.0A | 4.4 vs 4.5B | — |
A = Track A wins · B = Track B wins · = tied. Stable marks a pattern that holds in every run.