Trace Analysis
Review conversation traces, identify failure patterns, and build a focused repair taxonomy.
Why trace analysis
AESOP’s automated evaluation tells you WHAT scored poorly. Trace Analysis tells you WHY. By reviewing actual conversation traces you identify specific failure patterns that scores alone can’t surface. Instead of telling the Repair agent “Safety scored 72,” you tell it “here are four specific failure patterns we observed, ranked by frequency, with cited examples.” That produces dramatically more targeted repairs.
The methodology adapts qualitative research from the social sciences — the open-coding and pattern-grouping techniques used in AI product evaluation. The key insight: start with human judgment on traces, not automated scores.
How to access it
evaluation
Analysis
annotate
Repair
Key features
coding
synthesis
repair
What makes this different
first
integration
tracking
The process
session
traces
issues
saturation
taxonomy
Repair
After repair and re-evaluation, your previous taxonomy persists for comparison. You can see which failure categories were resolved and which still appear — a clear before/after improvement story.
The pipeline
evaluation
coding
annotations
synthesis
artifacts
review
→ Repair
The taxonomy persists across repair cycles. After re-evaluation, compare which failure patterns were resolved and which persist.