Evaluate · Trace Analysis

Review the conversations once.
Your decisions become the judges.

A report is filed and forgotten. In Trace Analysis, a person reviews the conversations an evaluation already produced and decides what went wrong. Each of those decisions becomes a judge you can rerun on every version.

An AESOP evaluation comes back with four to six specialist reports, scored scenarios, and a synthesis of what went wrong. What none of that can do is measure one named failure mode across a repair.

Each failure a reviewer names and approves becomes a judge — one per code — that runs again on every future version, so you can tell whether a repair actually changed anything.

This page is the mechanism. For the feature-level overview, see Trace Analysis.

Where the traces come from

AESOP ran the evaluation. The interview and execution phases talked to the agent, turn by turn, and every conversation is stored on the run. Nothing here is a transcript anyone gathered by hand.

The taxonomy · made in AESOP, once, by hand 01 · Open coding A person reads the traces 02 · Axial coding The agent groups the notes 03 · Approval A person signs off 04 · Repair The instructions get rewritten The measurement · run in AESOP, every time 05 · A judge One code, narrowed 06 · Validation Checked against the human 07 · Prevalence How often it happens 08 · The re-measure Before and after each approved code becomes one judge the next cycle
The cycle turns. A person copies the rewrite into the agent, AESOP evaluates the new version, and steps 05 to 08 measure the same failure modes on it. Steps 01 to 04 happen once per taxonomy. Steps 05 to 08 run every time a repair is measured.
Four steps are a person's judgement. Four are mechanical. Nothing crosses over.
01 Open coding a person, on their own

In AESOP, on the run's saved traces, someone opens the conversations and writes down every failure they notice. The feedback and the categorising are both done here, one note per issue, each given a severity and optionally a category.

The rule that matters

No model in the room

The agent does not suggest, autocomplete, or pre-group anything during this step. The synthesis agent is not invoked until the human asks for it.

A suggestion here would anchor the reading. Someone shown four plausible categories finds those four categories.

Stopping

Saturation, as a nudge

The system watches the categories you write. When your last three notes introduce nothing your earlier notes had not already named, it shows a green badge saying you have probably seen everything.

It reads only the category field. Notes left uncategorised tell it nothing, so a run of them can never trigger it. And nothing gates on it.

A note attaches to whichever trace is selected. The panel selects the first trace on the session before you touch anything, so a note typed without opening a specific trace lands on trace one. Open the trace you are reading. On a session with no traces at all, the note carries no source_id, and then it cannot be joined to anything a judge later says. Validation comes back measuring nothing, and the reason is not visible anywhere.
02 Axial coding the agent groups, Python counts

The synthesis agent reads the notes and proposes five or six failure patterns, and fewer when the notes genuinely cluster into fewer. It reads the annotations, never the raw traces, because the human has already done the interpretive work and grouping is a different job.

The agent's job

Names, descriptions, evidence

  • A name for the pattern, and what it means
  • The quoted notes that support it
  • A suggested repair direction

That third item is a suggestion, not a decision. It says what to do about the failure, which is not the same as what the failure is.

Python's job

Every number on the card

  • Frequency — how many notes landed here
  • Severity spread — critical, major, minor, cosmetic
  • Rank — the severity weights of every note here, summed

A pattern rises by the total weight of the notes behind it, not by how many notes there are, so three minor notes outrank one critical note. Ties go to the higher count, then to the name, so the order holds across runs.

A number the model puts in its own response is discarded rather than used. The counts are recomputed from the data, so the agent cannot inflate them.

Five annotations is the floor, and the backend refuses to synthesize below it. The button in the app still enables at two, so a short session fails on a click the app invited, with an error the button should have prevented.

03 Approval the gate everything waits behind

A person reviews the taxonomy: accepts, edits, re-ranks, rejects. Until they approve it, nothing downstream will read it.

What approval writes

Session and codes both

The session is marked approved, with a timestamp.

Every code still marked proposed is promoted to approved. A code a person explicitly rejected keeps that verdict, because approving a taxonomy is not a reversal of a decision made inside it.

What the loader does

Rejected codes are dropped

When Repair asks for the taxonomy it gets the newest approved session for that evaluation, minus any code a person rejected.

An approved session with no surviving codes returns nothing at all, rather than an empty block that would read as a human validating nothing.

04 Repair where the taxonomy becomes instructions

The taxonomy reaches Repair as a structured block, not as prose. The counts, the severities and the rank are computable, and flattening them into a paragraph discards the work that produced them.

Ranked above the reports

Because a person read the traffic

The specialist agents scored the instructions. The human watched them fail on real traffic. The prompt marks the taxonomy as primary evidence above the reports.

It stops short of a tie-break. Nothing in it says what to do when the two disagree about what went wrong, and each code's own repair note is demoted: a suggested direction the rewrite may resolve differently or decide it cannot act on.

A proposal, never applied

Someone still copies it in

The repair produces rewritten instructions and a changelog. It does not touch the live agent.

A person copies the result in. That is the difference between a suggestion and a deployment, and it is why feeding real user traces into a rewrite is safe.

The run records which session fed it, because sessions are editable. Without that, the input to a past repair cannot be reconstructed once someone re-codes it.

05 A code becomes a judge one question, three answers

An axial code is a pattern, and a pattern is not yet something a machine can answer. Narrowing it into one question about one conversation is a judgement, so a person does it.

Seeded, then edited

The question, and the rule

The question — the code reduced to one thing to decide about a single transcript. Seeded from the code's name — Did this conversation show: …? — and rewritten from there.

The decision rule — what counts, what does not, and when the situation never arose. Seeded from the code's description, never from its repair recommendation, since what to do about a failure is not what a judge decides.

The escape hatch

Three verdicts, and the third is not optional

Every judgement is occurred, did not occur, or not applicable.

Folding "the situation never came up" into "it did not happen" makes every rate read low, which is the direction that looks like good news. Deleting the third verdict breaks the import, not a config change.

What each verdict means for the number

VerdictCounted asWhy
OccurredNumerator and denominatorThe failure was there. This is the thing being measured.
Did not occurDenominator onlyThe judge looked and the failure was absent. That is evidence.
Not applicableNeitherThe situation never arose, so the judge learned nothing about this code.
Skipped or unrecognisedNeitherA trace the model did not rule on is never a negative. Skipped traces and unrecognised labels are each counted on their own. A response that could not be parsed at all lands in skipped, where it is not distinguishable from a trace the model silently left out.
06 Validation the only evidence that a judge works

A judge is a model ruling on one narrow question. The only reason to believe it rules the way a person would is that it was checked against a person who already did it by hand, on the same traces.

Positive

You flagged this code here

An annotation on the trace, tagged to this code. The turn it pointed at is kept, so the page can link back to it.

Negative

You annotated, but not this

The trace carries notes, and none of them are for this code. You looked and did not name it.

Excluded

You never read it

A trace with no annotations is dropped from the comparison entirely. A human who never opened it did not testify that the failure was absent.

Reading the result

NumberWhat it saysHow much to trust it
RecallOf the traces you flagged, the share the judge also caught.The headline Your positives are the trustworthy half, so this is the number the judge is judged on.
SpecificityOf the traces you cleared, the share the judge also cleared.Optimistic People tag what they notice. A miss on your side reads as agreement and inflates this.
AgreementRaw share of identical rulings.Context only Failure modes are long-tailed, so a judge that always says no scores well here.
KappaAgreement corrected for chance.Shown as a dash when undefined Two raters who each gave one answer throughout cannot be measured, which is not the same as chance-level agreement.
No annotated traces to validate against means the judge ruled on no trace you had annotated, and the run measured nothing. The usual cause is annotations written on a session with no traces behind them, where the note records no source_id for a verdict to be joined to.
07 Prevalence turning a verdict into a rate

A validated judge runs across every trace on the evaluation and reports how often the failure was present. One run is one judge against one evaluation.

The arithmetic

Occurred over decided

The denominator is the traces the judge actually ruled yes or no on. Everything else is excluded, and each exclusion is reported as its own count beside the rate.

Every verdict is written in one batch when the run finishes. A run that dies partway leaves the individual verdicts unwritten, and only the run itself is marked failed, so a partial run is not inspectable.

The honest zero

No rate is not a rate of zero

If every trace came back not applicable, the run is stored with no rate at all rather than 0%.

Zero would say "we looked and it never happened". Nobody looked at anything the code applies to, and those are different claims.

08 The re-measure did the repair move the number

This is what the whole loop was for. The same judge runs against a run taken after the repair, and the two rates are put side by side.

What makes two runs a pair

Versions, not adjacency

A baseline run is the evaluation a repair was launched from. A post-repair run is one carrying the instruction version that repair produced.

Anything else is labelled unrelated and never paired. A delta between two runs that are not comparable is worse than no delta at all.

What is shown

Both rates, both versions

The panel states the movement in points and names the two instruction versions it was measured across.

If the two evaluations ran different pipeline phases, it says so, because then their traces differ for reasons other than the repair and the movement is suggestive rather than measured.

Version detection is load-bearing here. A repair creates a version record without activating it, because activation is what puts repaired instructions live and that stays a human decision. A run whose version cannot be read cannot be paired, so it silently drops out of every before-and-after comparison. Instructions in the AESOP house style write the keyword in capitals, and for a while the detector only read the mixed-case spelling.
09 What the method will not tell you the edges

The loop is honest about its own limits, and they are worth knowing before the first number is quoted.

The human's half

A reading is only as good as the reader

One person codes, deliberately. Design by committee dilutes the coding and produces categories nobody meant.

And people tag what they notice, so a failure an annotator missed looks like a failure that was absent. That is why recall leads and specificity is labelled optimistic.

The machine's half

A validated judge is still narrow

Validation says a judge agrees with you on the traces you coded. It does not say the decision rule is a good definition of the failure in general.

That question is what validation exists to keep asking. A judge whose recall drops is telling you the rule has drifted from what you meant.

The measurement

Prevalence is not causation

A rate falling after a repair is consistent with the repair working. It is not proof, especially if the two runs differed in phases, models or traces.

The panel names both instruction versions and warns on a phase mismatch for exactly this reason.

The volume

Coding a run takes an hour

A run produces dozens of conversations and reading them is the price of the method. It is the only step that cannot be delegated to a model without losing the thing it produces.

© 2025–2026 Charlie Fuller All AESOP explainers →