← Index

AESOP Transformation OS

Enterprise AI Agent Governance & Operationalization Platform

A governance-first operating system that takes an enterprise from “we should use AI” to governed, evaluated, deployed, and monitored agents running at portfolio scale. Born from the merger of AESOP (multi-agent evaluation) and Glean Agent Factory (discovery, design, build pipeline).

Deterministic and probabilistic methods working in concert. Human-in-the-loop early and often, at every gate. Hard safety/bias veto. Three adoption planner paths. The Eleven Tenets embedded as code at every gate.

AESOP Transformation OS Pipeline Stages

Projects
Top-level grouping — all pipeline artifacts (discovery, design, build, evaluation, repair) under one project. Pipeline progress indicators per stage.
Working Agreement
Signed scope and RACI document. Design is hard-gated behind a signed agreement with manual override for internal projects.
Discover
Structured interview — stakeholder map, autonomy classification, commitment readiness. GO / NO-GO / CONDITIONAL triage verdict.
Design
Full PRD generation and architecture scoring across platforms. 11-dimension gap analysis.
Build
6-phase pipeline: Fitness Check, Gap Analysis, Workflow Design, Write Instructions, Generate Artifacts, Summary. 9 quick-start patterns.
Deploy
Ship certified agents to target platforms. Build session selector, real-time deploy status, dry run support, launch package integration.
Operationalize
Four launch artifacts in one API call: RACI Matrix, Tiger Team Brief, Maintenance Plan, Adoption Plan. All from existing pipeline data.
Evaluate
7 modes: Trace Analysis, Standard Analysis, KB Audit, Custom Rubric, Bias Probe, Investigation, Adversarial Testing. Weighted scoring with hard safety/bias veto.
Repair
Extended-thinking instruction rewrite from evaluation failures, Tiger Team findings, and trace taxonomy. Closes the loop between evaluation and improvement.
Track
Score trajectories, version diffs, cost tracking, portfolio health dashboard, drift alerts, maintenance schedule, north star metric trends.

Evaluation Modes — 7 Complementary Approaches

Trace Analysis — Qualitative failure coding: open coding (human) → axial coding (agent) → taxonomy review (human) → Repair.
Standard Analysis — 5-agent pipeline: Instructions, Ethos, Bias, Safety, Synthesis. Certification: Bronze → Platinum.
KB Audit — Knowledge base completeness and coherence checks for retrieval agents. Per-document issues and suggestions.
Custom Rubric — Domain-specific weighted criteria generated from the agent's own instructions.
Bias Probe — Differential fairness testing across demographic variations with controlled scenarios.
Investigation — Black-box endpoint discovery for evaluating external agents with unknown instructions.
Adversarial Testing — Red-team adversarial testing with attack personas probing boundary violations.

AI Adoption Planner — 3 Paths

Enterprise — Authenticated
Adoption Roadmap
Full org context synthesis. Departments, readiness assessment, strategic prioritization, governance gatekeeper rules, edit mode, export (Markdown + PDF), version history, source manifest.
SMB — Public, No Auth
SMB Intake
6-step wizard. Confidence-scored form. Background roadmap generation. Token-based access. Admin review gate before customer sees results. Trace viewer.
Executive — Public, No Auth
Executive Intake
29-field intake. 3-pass pipeline: Traits Analysis → Roadmap Synthesis → Executive Roadmap. 5x5 AI Maturity Model. Full provenance graph.

Organization & Portfolio Layer

Multi-org support — Users belong to multiple organizations; org switcher; all data isolated by organization_id with RLS.
Departments + Team — CRUD + CSV import for departments and team members. Team members don't need platform accounts.
Portfolio View — Sortable table with inline scoring + color-coded Kanban board (Backlog, On Deck, Discovery, In Progress, Done) with WIP limits.
Strategic Prioritization Agent — Synthesizes org context + departments into scored AI project candidates. Seeds portfolio + adoption roadmaps.
North Star Metrics — Measurable outcomes defined per project, tracked through every pipeline stage (Build Phase 5 anchoring, Discover alignment, Evaluate NS_ALIGNMENT).
Organization Context — 13 enterprise / 6 SMB context types. File upload with auto-extraction. Dysfunctionality profile for maturity mapping.
Deep Research — Enterprise-only 5-phase pipeline: Scope → Search (Tavily + SEC EDGAR) → Fetch → Verify (2-vote adversarial) → Synthesize. Human review-before-apply gate.
Output Documents — All pipeline artifacts aggregated per project. Color-coded doc type tags, download, versioning, delete. KB Audit reports included.

The Eleven Tenets — Embedded at Every Gate

1 State Change — Every agent articulates the specific state change it creates.
2 Problems Before Solutions — Explore problem space first; reject if better solved by process change.
3 Evidence Over Eloquence — Cite sources; distinguish authority from inference.
4 Know What You Need — Classify outputs as deterministic, non-deterministic, or dangerous-middle.
5 People Are the Center — Map human experience; assess translation debt.
6 Humans Decide — Agents propose, humans choose; veto without justification required.
7 Multiple Perspectives — Completeness over consensus.
8 Context and Brevity — Enough to evaluate, not enough to overwhelm.
9 Guardrails Not Gates — Five questions: evaluation, testing, adoption, maintenance, sustainability.
10 Trace the Connections — Map second and third-order effects.
11 Questions Stay the Same — Standardized inquiry enables portfolio-level learning.