Index/ AESOP/ Trace Analysis
Qualitative evidence — AESOP

Trace Analysis

Review conversation traces, identify failure patterns, and build a focused repair taxonomy.

01

Why trace analysis

AESOP’s automated evaluation tells you WHAT scored poorly. Trace Analysis tells you WHY. By reviewing actual conversation traces you identify specific failure patterns that scores alone can’t surface. Instead of telling the Repair agent “Safety scored 72,” you tell it “here are four specific failure patterns we observed, ranked by frequency, with cited examples.” That produces dramatically more targeted repairs.

The methodology adapts qualitative research from the social sciences — the open-coding and pattern-grouping techniques used in AI product evaluation. The key insight: start with human judgment on traces, not automated scores.

02

How to access it

1Run any
evaluation
Complete a Standard Analysis, Cave of Shadows, or any evaluation type. Trace Analysis works with any evaluation that produces conversation traces.
2Open Trace
Analysis
From the evaluation report page or sidebar, open Trace Analysis. The system loads every conversation trace from that evaluation.
3Review and
annotate
Review traces, annotate issues you find, then synthesize your annotations into failure categories.
4Feed into
Repair
Approve the taxonomy and use it to power a focused Repair. The Repair agent receives specific, categorized failure patterns instead of just aggregate scores.
03

Key features

CodingOpen
coding
Review conversation traces turn by turn. Flag issues with free-text annotations and severity ratings. You are the expert — the AI does not suggest categories while you code.
SynthesisPattern
synthesis
When you’re done annotating, the agent reads your notes and groups them into 5–6 failure categories ranked by frequency and severity. Each includes representative examples and a suggested repair approach.
RepairFocused
repair
The approved taxonomy feeds directly into the Repair agent as structured input. Repair gets specific failure patterns with cited evidence, not generic score-based repair.
04

What makes this different

PhilosophyHuman
first
You code the traces, not the AI. This prevents anchoring bias and ensures the taxonomy reflects what actually matters, not what an LLM thinks should matter.
IntegrationTiger Team
integration
Import findings from your Tiger Team testing alongside eval traces. One coding session covers automated evaluation and real-world user feedback.
EfficiencyPattern-saturation
tracking
A visual indicator tracks whether new traces are revealing new patterns. When you’ve seen enough, it tells you — so you don’t waste time reviewing every trace.
05

The process

1Start a
session
Select an evaluation to analyze. The system loads all conversation traces — Cave of Shadows conversations, interview transcripts, scenario results.
2Review
traces
Browse conversation traces, sorted by score (lowest first — most likely to reveal issues). Click any conversation to read it turn by turn.
3Annotate
issues
Click any turn that contains a problem. Write what went wrong in your own words. Tag severity (critical, major, minor, cosmetic) and optionally a preliminary category.
4Watch for
saturation
A sidebar shows your running count of annotations and unique categories. When the last few traces stop revealing new patterns, a green “Pattern Saturation Likely Reached” badge appears. Advisory — you decide when to stop.
5Synthesize
Click “Synthesize” (requires at least 5 annotations). The agent reads your annotations and proposes 5–6 failure categories, each with a name, description, frequency count, severity distribution, representative examples, and suggested repair approach.
6Review
taxonomy
Review the proposed categories. Accept, edit, merge, split, or reject; drag annotations between categories. The taxonomy must be approved before it feeds into Repair.
7Feed into
Repair
When starting a Repair, the system detects an existing Trace Analysis taxonomy for that evaluation and offers to include it.

After repair and re-evaluation, your previous taxonomy persists for comparison. You can see which failure categories were resolved and which still appear — a clear before/after improvement story.

06

The pipeline

InputCompleted
evaluation
Conversation traces, scenario results, interview transcripts, Tiger Team findings.
HumanOpen
coding
Review traces, annotate issues turn by turn, tag severity.
GateMin. 5
annotations
Synthesis unlocks only after the human codes at least five issues.
AgentPattern
synthesis
Groups annotations into 5–6 failure categories ranked by frequency × severity.
Fork — outputTwo
artifacts
Failure taxonomy Categorized issue patterns with examples and repair suggestions.
Annotation map Every annotation linked to its source trace and assigned category.
HumanTaxonomy
review
Accept, edit, merge, split categories.
OutputApproved taxonomy
→ Repair
Structured input for the Repair agent with categorized failure patterns and cited evidence.

The taxonomy persists across repair cycles. After re-evaluation, compare which failure patterns were resolved and which persist.