← Index

The Pipeline

Every project passes through every stage. Each stage produces artifacts that feed the next.

Stages & Gates
Projects contain all pipeline artifacts. Stages are sequential. Gates enforce human approval before continuing.
CTX
Foundation
Organization Context
The grounding layer for everything below it. Captured context — authority & decision rights, coordination & dependencies, tacit knowledge, exposure posture, systemic blockers, plus departments and uploaded org documents — feeds the Delegation Readiness assessment, sharpens prioritization, and informs discovery. The sharper this context, the sharper every downstream judgment.
authority & decision rights coordination map systemic blockers departments org documents
ORG
Organization Layer
Portfolio & Delegation Readiness
The layer above individual projects. The Portfolio coordinates every project across departments and ranks them by strategic score; Delegation Readiness assesses whether the organization can responsibly scale agent autonomy and turns gaps into tracked action items — which can be promoted straight into the portfolio.
portfolio board & table strategic prioritization readiness verdict org action items
PRJ
Container
Projects
Top-level initiative grouping. A project wraps every artifact produced across the entire pipeline — discovery sessions, agreements, builds, evaluations, repairs — under a single entity for portfolio-level tracking.
project_id propagation status tracking portfolio analytics
WA
Pre-Engagement Gate
Working Agreement
A signed document establishing scope, responsibilities, and response commitments. Rendered from structured discovery data into a static template. Design is gated behind a signed agreement (with manual override for internal projects).
signed agreement RACI SLAs
Signature Required
01
Stage 1
Discover
Structured problem interview that captures stakeholder map, autonomy classification, and commitment readiness. Produces an Opportunity Brief with a triage verdict (GO / NO_GO / CONDITIONAL). Platform-agnostic — not locked to any implementation.
opportunity brief stakeholder map triage verdict autonomy level
02
Stage 2
Design
Claude-powered PRD generation with multi-platform architecture recommendation. An ArchitectureAgent evaluates the PRD against platform capabilities (Glean, Claude API, OpenAI, Generic) and produces a gap analysis across 11 dimensions.
PRD architecture recommendation gap analysis platform scores
03
Stage 3
Build
Six-phase instruction writing pipeline with a human-in-the-loop checkpoint after gap analysis. Phases: Fitness Check, Gap Analysis (HITL), Workflow Design, Write Instructions, Generate Artifacts, Summary. Real-time SSE streaming to the UI.
system instructions platform artifacts state change hypothesis translation debt assessment
HITL Approval at Phase 1
3.5
Stage 3.5
Deploy
Ship certified agents to target platforms. Build session selector, real-time deploy status, dry run support, launch package integration. Platform connectors (8+2): Anthropic, Glean, OpenAI, Google, NemoClaw, REST, Base, Factory, WebSocket, Managed Session.
live deployment session selector deploy status dashboard platform connectors
3.6
Stage 3.6
Operationalize
Generates a Launch Package — four organizational artifacts from pipeline data already captured during discovery and build. No new data collection required; a generation pass over existing outputs.
RACI matrix tiger team brief maintenance plan adoption plan
04
Stage 4
Evaluate
Multi-agent scoring across 7 evaluation modes: Standard Analysis, Cave of Shadows (red-team), Custom Rubric, Bias Probe, Investigation, KB Audit, and Trace Analysis.
The name "Cave of Shadows" draws from Plato's Allegory of the Cave, where prisoners mistake shadows on the wall for reality. In the same way, AI systems can project confident, convincing responses that mask underlying flaws — bias, hallucination, boundary violations. The Adversarial Red-Team capability drags those shadows into the light, testing whether an AI holds up when confronted with adversarial personas and edge-case scenarios designed to expose what lurks beneath the surface.
scored reports certification level risk strategy attack scenarios
4.5
Stage 4.5
Regulatory Mapping
Upload your regulatory templates (ISO 13485, risk registers, change control forms). An agent identifies fillable fields, maps them to AESOP pipeline evidence, and produces completed documents for QMS filing. AESOP is evidence infrastructure, not a regulatory authority.
completed templates field mappings evidence traceability
05
Stage 5
Repair
Uses extended thinking to deeply analyze evaluation failures and rewrite system instructions. Takes evaluation context, scenario results, trace analysis taxonomy, and optionally Tiger Team findings as input. Closes the loop between evaluation, stakeholder feedback, and improvement.
rewritten instructions change rationale diff summary
06
Stage 6
Track
Score trajectory visualization, version history with diffs, cost tracking per agent and in aggregate, portfolio-level analytics, and maintenance schedule dashboard with drift alerts.
score trajectories version diffs cost analytics drift alerts
Closed-Loop Improvement
The Evaluate-Repair-Track cycle runs continuously
Evaluate
Repair
Re-Evaluate
Track Delta
Tiger Team findings feed back into Repair
Launch Package
Stakeholder Review
Repair
The Eleven Tenets
The decision framework embedded in every pipeline stage. Not just documentation — methodology as code.
1
State Change
2
Problems Before Solutions
3
Evidence Over Eloquence
4
Know What You Need
5
People Are the Center
6
Humans Decide
7
Multiple Perspectives
8
Context and Brevity
9
Guardrails Not Gates
10
Trace the Connections
11
The Questions Stay the Same