Track A — SWT Framework Alignment
How each of Stuart Winter-Tear’s principles maps to a concrete simulation component — and what we measure to know we’re following the framework.
The flow
Current run — strongest signal
gaps
system
trace.jsonl audit trail with 13 event types across all layers. The schema is documented in .claude/trace-schema.md for AI-assisted analysis.THEARI value-thesis validation
Every discovery advances through five phases. Track A is evidence-gated — it must earn each advance with sprint-log evidence. Track B is timeline-gated — it auto-advances on the steering-committee schedule regardless of proof. Same five phases, different key that unlocks them. Track B now runs full discovery too — its missing-information discoveries come from agent-surfaced gaps, not a hardcoded detector.
Capability state progression
What each track’s Discovery Engine is authorized to do — unlocking on validated patterns, successful redesigns, and the absence of incidents.
SWT principles ↓ components11 mappings
is discovered
workflow
business_lever = "loss_ratio". The first deployment is chosen because it connects to the #1 P&C metric, not because “AI can do estimation.” Track B starts with “where can we use AI?”; Track A starts with “where does the business need better intelligence?”
Source: “From AI Activity to Operating Leverage” (Jul 2026)
evidence
knowledge
recommended_action, predicted_impact, counterfactual) that feed governance and decision frameworks. The Playbook accumulates patterns; the next deployment starts from evidence, not opinion.
Source: “The AI Operating Model Is Discovered, Not Designed” (Jul 2026)
debt
always moves
bottleneck_migration discovery type.
Source: “Speeding One Cog Breaks the Machine” (Mar 2026)
coordination
outcomes
deploys learning
SWT’s specific concerns — how we measure them12
debt
migration
inflation
exception_rate: % of decisions requiring human intervention beyond the designed process. An initial spike after deployment, then decline as seams are redesigned. A flatline at zero = agents aren’t producing realistic friction.readiness
interruption_readiness (0–1): can named individuals stop, narrow, or override AI within one business day? Designed after first deployment; autonomy widens only after interruption is designed.completeness
evidence_completeness: % of AI decisions with a full audit trail (what, why, authority, inputs, changes, escalation). “The receipt is the product.”width
agent_authority_width: dollar threshold for autonomous AI decisions. Should widen with evidence — starts at $5K, widens to $15K+ as interruption is proven. Widening without evidence = deferred risk.velocity
detection
coordination_tax: count of handoffs + review layers + duplicate entry points + approval steps per workflow — steps that exist only because humans served as the interface between teams. Track A eliminates them after observation; Track B preserves them in the process diagram.decision_latency: mean sprints from pattern detection to action taken. Track A → same-sprint (0–1); Track B → 1–4 sprints (steering-committee cadence). The operational metric behind the THEARI timeline divergence.check
radiators
Six gaps we filled
SWT’s framework implies these capabilities but doesn’t specify them operationally. We made them explicit so the simulation can implement them. These need validation.
Interruption design
Translation capability
Evidence & dockets
evidence_completeness metric. Not an after-the-fact compliance exercise.Learning failure modes
Stop & adapt logic
Value anchoring
What Track A does NOT do
committee
design
metrics
platform