Proof framework

Measurement & Validation

How we know the simulation captures reality, how we measure transformation, and how we prove the results are auditable.

01

Capture architecture7 artifacts

decisions.jsonl
Every agent decision — role, step, reasoning, concerns, AI interaction, emotional state, translation incidents. The raw ground truth.
metrics.csv
22 metrics × 2 tracks × 16 sprints. A time series for every transformation dimension.
state.json
Full playbook, governance rules, decision frameworks, trust values, steering-committee decisions, phases.
forensic_record.json
Master document — temporal timeline, causal chains, comparative analysis, alignment check, operational metadata.
comprehensive_report.json
All analyses merged: agent breakdowns, trust timeline, bottleneck log, claim traces, market projection, divergence, psych differentiation, fidelity.
evidence_citations.json
29 real-world sources grounding every parameter. 34% official (SEC filings, BLS), 17% press, 14% aggregator, 21% inferred, 14% assumption.
audit_trail.json +
audit_verification.json
Every derived number traceable to source data. 8/8 outputs auditable (100%). Verification confirms evidence, derivation, and source for each output.
02

Six transformation dimensions

D1Financial
performance
6 metrics. Combined ratio (87.3% baseline), loss ratio (64%), expense ratio (23.3%), cost per claim, premium growth, ROI. The market model projects operational metrics → EPS, stock price, market cap.
D2Operational
performance
12 metrics. Cycle time, exception rate, first-pass accuracy, supplement rate, translation-debt index, handoff-failure rate, evidence completeness, interruption readiness, agent authority width, and four SWT-specific operational metrics.
D3Organizational
health
9 metrics. Employee AI trust, trust polarization, claims-adjuster attrition, learning-loop closure rate, pattern-reuse rate, informal control points surfaced, playbook growth, change-exhaustion trend.
D4Customer
impact
4 metrics. Customer retention after claim (proxy), CSAT (directional), claim error rate, time to first contact.
D5Competitive
position
4 metrics. Market share, competitive absorption speed, pricing accuracy, Robinsons segment growth. Benchmarked against GEICO (91.2 CR), Allstate (86.6), industry (92.0).
D6Governance
& risk
6 metrics. Regulatory findings, AI bias incidents, audit-trail completeness, intervention response time, model-drift events, stop-button testability.
03

Audit chain

Every output is traceable end to end. Nothing asserted without origin.

Real-world data Evidence registry Simulation parameter Agent decision Metric Projection Audit trail

Source data → derivation → confidence → evidence IDs. Every output traceable to its origin.

04

Validity gates must all pass

PASSG1 · Psych
differentiation
DeepSeek diagnostic: Diana uses AI 26%, Maria 81%. Different agents make different decisions on the same claims. Passed.
PASSG2 · Translation debt
surfaces
DeepSeek baseline: 35 translation incidents (32.7% of decisions). Handoff friction emerges organically. AI contamination 0%. Passed.
PASSG3 · Bottleneck
migrates
Mock 16-sprint: 35 bottleneck detections across workflow steps. The constraint moves after deployment. To verify with DeepSeek. Passed (mock).
PASSG4 · Track B produces
some improvement
Round 4 sweep: after the fairness fixes, B’s translation debt collapses from ~23% to ~5% in the codifiable regime — a tie with A, not an 18-point loss. B now produces real improvement where its a-priori design is accurate. Passed.
PASSG5 · Divergence is
traceable
Round 3 traced B’s loss to three mechanisms (wrong seams targeted, wash-off regression, one-shot rollout); Round 4 confirmed the fix in the trace. Each divergence point traces to a decision chain. Passed.
PASSG6 · No zero-variance
metrics
Mock 16-sprint: all metrics show variance across sprints; no flatlines. DeepSeek baseline confirms. Passed.
PASSG7 · Fidelity auditor
functional
Mock 16-sprint: 19 findings across 2 categories. Anti-caricature checks operational. Passed.
05

What proof looks like

The simulation
surprises us
Novelty in the novelty register is the best signal. If every output confirms what we already believe, we built an echo chamber.
Practitioners recognize
the behavior
If a claims VP reads Diana’s decision log and says “I’ve managed three of her,” the model captures something real. Face validity from domain experts.
Results are robust,
not fragile
Run 5 times with different seeds. Specific discoveries vary; structural patterns hold. Small parameter changes don’t flip the results.
Track B isn’t
a straw man
Track B produces some improvement; the steering committee makes some good decisions. If Track B is a punching bag, the experiment is propaganda.
The counterfactual
is honest
Can we construct conditions where Track B outperforms? If political friction is low, Diana cooperates, and the existing process is clean, traditional change management might win.