Proof framework
Measurement & Validation
How we know the simulation captures reality, how we measure transformation, and how we prove the results are auditable.
01
Capture architecture7 artifacts
decisions.jsonl
Every agent decision — role, step, reasoning, concerns, AI interaction, emotional state, translation incidents. The raw ground truth.
metrics.csv
22 metrics × 2 tracks × 16 sprints. A time series for every transformation dimension.
state.json
Full playbook, governance rules, decision frameworks, trust values, steering-committee decisions, phases.
forensic_record.json
Master document — temporal timeline, causal chains, comparative analysis, alignment check, operational metadata.
comprehensive_report.json
All analyses merged: agent breakdowns, trust timeline, bottleneck log, claim traces, market projection, divergence, psych differentiation, fidelity.
evidence_citations.json
29 real-world sources grounding every parameter. 34% official (SEC filings, BLS), 17% press, 14% aggregator, 21% inferred, 14% assumption.
audit_trail.json +
audit_verification.json
audit_verification.json
Every derived number traceable to source data. 8/8 outputs auditable (100%). Verification confirms evidence, derivation, and source for each output.
02
Six transformation dimensions
D1Financial
performance
performance
6 metrics. Combined ratio (87.3% baseline), loss ratio (64%), expense ratio (23.3%), cost per claim, premium growth, ROI. The market model projects operational metrics → EPS, stock price, market cap.
D2Operational
performance
performance
12 metrics. Cycle time, exception rate, first-pass accuracy, supplement rate, translation-debt index, handoff-failure rate, evidence completeness, interruption readiness, agent authority width, and four SWT-specific operational metrics.
D3Organizational
health
health
9 metrics. Employee AI trust, trust polarization, claims-adjuster attrition, learning-loop closure rate, pattern-reuse rate, informal control points surfaced, playbook growth, change-exhaustion trend.
D4Customer
impact
impact
4 metrics. Customer retention after claim (proxy), CSAT (directional), claim error rate, time to first contact.
D5Competitive
position
position
4 metrics. Market share, competitive absorption speed, pricing accuracy, Robinsons segment growth. Benchmarked against GEICO (91.2 CR), Allstate (86.6), industry (92.0).
D6Governance
& risk
& risk
6 metrics. Regulatory findings, AI bias incidents, audit-trail completeness, intervention response time, model-drift events, stop-button testability.
03
Audit chain
Every output is traceable end to end. Nothing asserted without origin.
Real-world data → Evidence registry → Simulation parameter → Agent decision → Metric → Projection → Audit trail
Source data → derivation → confidence → evidence IDs. Every output traceable to its origin.
04
Validity gates must all pass
PASSG1 · Psych
differentiation
differentiation
DeepSeek diagnostic: Diana uses AI 26%, Maria 81%. Different agents make different decisions on the same claims. Passed.
PASSG2 · Translation debt
surfaces
surfaces
DeepSeek baseline: 35 translation incidents (32.7% of decisions). Handoff friction emerges organically. AI contamination 0%. Passed.
PASSG3 · Bottleneck
migrates
migrates
Mock 16-sprint: 35 bottleneck detections across workflow steps. The constraint moves after deployment. To verify with DeepSeek. Passed (mock).
PASSG4 · Track B produces
some improvement
some improvement
Round 4 sweep: after the fairness fixes, B’s translation debt collapses from ~23% to ~5% in the codifiable regime — a tie with A, not an 18-point loss. B now produces real improvement where its a-priori design is accurate. Passed.
PASSG5 · Divergence is
traceable
traceable
Round 3 traced B’s loss to three mechanisms (wrong seams targeted, wash-off regression, one-shot rollout); Round 4 confirmed the fix in the trace. Each divergence point traces to a decision chain. Passed.
PASSG6 · No zero-variance
metrics
metrics
Mock 16-sprint: all metrics show variance across sprints; no flatlines. DeepSeek baseline confirms. Passed.
PASSG7 · Fidelity auditor
functional
functional
Mock 16-sprint: 19 findings across 2 categories. Anti-caricature checks operational. Passed.
05
What proof looks like
The simulation
surprises us
surprises us
Novelty in the novelty register is the best signal. If every output confirms what we already believe, we built an echo chamber.
Practitioners recognize
the behavior
the behavior
If a claims VP reads Diana’s decision log and says “I’ve managed three of her,” the model captures something real. Face validity from domain experts.
Results are robust,
not fragile
not fragile
Run 5 times with different seeds. Specific discoveries vary; structural patterns hold. Small parameter changes don’t flip the results.
Track B isn’t
a straw man
a straw man
Track B produces some improvement; the steering committee makes some good decisions. If Track B is a punching bag, the experiment is propaganda.
The counterfactual
is honest
is honest
Can we construct conditions where Track B outperforms? If political friction is low, Diana cooperates, and the existing process is clean, traditional change management might win.