← Index
Qualitative Evidence

Trace Analysis

Review conversation traces, identify failure patterns, and create a focused repair taxonomy

Why Trace Analysis
AESOP's automated evaluation tells you WHAT scored poorly. Trace Analysis tells you WHY. By reviewing actual conversation traces from your evaluation, you identify specific failure patterns that scores alone can't surface. Instead of telling the Repair agent "Safety scored 72," you tell it "here are 4 specific failure patterns we observed, ranked by frequency, with cited examples from actual conversations." This produces dramatically more targeted repairs.

The methodology is adapted from qualitative research methods used in social science -- specifically the open coding and pattern grouping techniques described by Hamel Husain and Shreya Shankar for AI product evaluation. The key insight: start with human judgment on traces, not automated scores.
How to Access
1
Run Any Evaluation
Complete a Standard Analysis, Cave of Shadows, or any other evaluation type. Trace Analysis works with any evaluation that produces conversation traces.
2
Open Trace Analysis
From the evaluation report page or sidebar, open Trace Analysis. The system loads all conversation traces from that evaluation.
3
Review and Annotate
Review traces, annotate issues you find, then synthesize your annotations into failure categories.
4
Feed into Repair
Approve the taxonomy and use it to power a focused Repair. The Repair agent receives specific, categorized failure patterns instead of just aggregate scores.
Key Features
Coding
Open Coding
Review conversation traces turn by turn. Flag issues with free-text annotations and severity ratings. You are the expert -- the AI does not suggest categories while you code.
Synthesis
Pattern Synthesis
When you're done annotating, the agent reads your notes and groups them into 5-6 failure categories ranked by frequency and severity. Each category includes representative examples and a suggested repair approach.
Repair
Focused Repair
The approved taxonomy feeds directly into the Repair agent as structured input. Instead of generic score-based repair, Repair gets specific failure patterns with cited evidence.
What Makes This Different
Philosophy
Human First
You code the traces, not the AI. This prevents anchoring bias and ensures the taxonomy reflects what actually matters, not what an LLM thinks should matter.
Integration
Tiger Team Integration
Import findings from your Tiger Team testing alongside eval traces. One coding session covers both automated evaluation and real-world user feedback.
Efficiency
Pattern Saturation Tracking
A visual indicator tracks whether new traces are revealing new patterns. When you've seen enough, it tells you -- so you don't waste time reviewing every trace.
Step-by-Step Walkthrough
1
Start a Session
Select an evaluation to analyze. The system loads all conversation traces -- Cave of Shadows conversations, interview transcripts, scenario results -- from that evaluation.
2
Review Traces
Browse conversation traces. Traces are sorted by score (lowest first -- most likely to reveal issues). Click any conversation to read it turn by turn.
3
Annotate Issues
Click any turn that contains a problem. Write what went wrong in your own words. Tag the severity (critical, major, minor, cosmetic). Optionally tag with a preliminary category.
4
Watch for Pattern Saturation
A sidebar shows your running count of annotations and unique categories. When the last few traces stop revealing new patterns, a green "Pattern Saturation Likely Reached" badge appears. This is advisory -- you decide when to stop.
5
Synthesize
Click "Synthesize" (requires at least 5 annotations). The agent reads all your annotations and proposes 5-6 failure categories. Each category has a name, description, frequency count, severity distribution, representative examples, and suggested repair approach.
6
Review Taxonomy
Review the proposed categories. Accept, edit, merge, split, or reject categories. Drag annotations between categories. The taxonomy must be approved before it can feed into Repair.
7
Feed into Repair
When starting a Repair, the system detects if a Trace Analysis taxonomy exists for that evaluation and offers to include it. The Repair agent receives specific, categorized failure patterns instead of just aggregate scores.
After repair and re-evaluation, your previous taxonomy persists for comparison. You can see which failure categories were resolved and which still appear -- creating a clear before/after improvement story.
Trace Analysis Pipeline
Input
Source
Completed Evaluation
Conversation traces, scenario results, interview transcripts, Tiger Team findings
Human
Open Coding
Review traces, annotate issues turn by turn, tag severity
Min 5 Annotations
Agent
Pattern Synthesis
Groups annotations into 5-6 failure categories ranked by frequency x severity
Output
Artifact
Failure Taxonomy
Categorized issue patterns with examples and repair suggestions
Artifact
Annotation Map
Every annotation linked to its source trace and assigned category
Human
Taxonomy Review
Accept, edit, merge, split categories
Output
Approved Taxonomy → Repair
Structured input for the Repair agent with categorized failure patterns and cited evidence
The taxonomy persists across repair cycles. After re-evaluation, compare which failure patterns were resolved and which persist.