A report is filed and forgotten. In Trace Analysis, a person reviews the conversations an evaluation already produced and decides what went wrong. Each of those decisions becomes a judge you can rerun on every version.
An AESOP evaluation comes back with four to six specialist reports, scored scenarios, and a synthesis of what went wrong. What none of that can do is measure one named failure mode across a repair.
Each failure a reviewer names and approves becomes a judge — one per code — that runs again on every future version, so you can tell whether a repair actually changed anything.
This page is the mechanism. For the feature-level overview, see Trace Analysis.
AESOP ran the evaluation. The interview and execution phases talked to the agent, turn by turn, and every conversation is stored on the run. Nothing here is a transcript anyone gathered by hand.
In AESOP, on the run's saved traces, someone opens the conversations and writes down every failure they notice. The feedback and the categorising are both done here, one note per issue, each given a severity and optionally a category.
The agent does not suggest, autocomplete, or pre-group anything during this step. The synthesis agent is not invoked until the human asks for it.
A suggestion here would anchor the reading. Someone shown four plausible categories finds those four categories.
The system watches the categories you write. When your last three notes introduce nothing your earlier notes had not already named, it shows a green badge saying you have probably seen everything.
It reads only the category field. Notes left uncategorised tell it nothing, so a run of them can never trigger it. And nothing gates on it.
source_id, and then it cannot be joined to anything a judge later says. Validation comes back measuring nothing, and the reason is not visible anywhere.The synthesis agent reads the notes and proposes five or six failure patterns, and fewer when the notes genuinely cluster into fewer. It reads the annotations, never the raw traces, because the human has already done the interpretive work and grouping is a different job.
That third item is a suggestion, not a decision. It says what to do about the failure, which is not the same as what the failure is.
A pattern rises by the total weight of the notes behind it, not by how many notes there are, so three minor notes outrank one critical note. Ties go to the higher count, then to the name, so the order holds across runs.
A number the model puts in its own response is discarded rather than used. The counts are recomputed from the data, so the agent cannot inflate them.
Five annotations is the floor, and the backend refuses to synthesize below it. The button in the app still enables at two, so a short session fails on a click the app invited, with an error the button should have prevented.
A person reviews the taxonomy: accepts, edits, re-ranks, rejects. Until they approve it, nothing downstream will read it.
The session is marked approved, with a timestamp.
Every code still marked proposed is promoted to approved. A code a person explicitly rejected keeps that verdict, because approving a taxonomy is not a reversal of a decision made inside it.
When Repair asks for the taxonomy it gets the newest approved session for that evaluation, minus any code a person rejected.
An approved session with no surviving codes returns nothing at all, rather than an empty block that would read as a human validating nothing.
The taxonomy reaches Repair as a structured block, not as prose. The counts, the severities and the rank are computable, and flattening them into a paragraph discards the work that produced them.
The specialist agents scored the instructions. The human watched them fail on real traffic. The prompt marks the taxonomy as primary evidence above the reports.
It stops short of a tie-break. Nothing in it says what to do when the two disagree about what went wrong, and each code's own repair note is demoted: a suggested direction the rewrite may resolve differently or decide it cannot act on.
The repair produces rewritten instructions and a changelog. It does not touch the live agent.
A person copies the result in. That is the difference between a suggestion and a deployment, and it is why feeding real user traces into a rewrite is safe.
The run records which session fed it, because sessions are editable. Without that, the input to a past repair cannot be reconstructed once someone re-codes it.
An axial code is a pattern, and a pattern is not yet something a machine can answer. Narrowing it into one question about one conversation is a judgement, so a person does it.
The question — the code reduced to one thing to decide about a single transcript. Seeded from the code's name — Did this conversation show: …? — and rewritten from there.
The decision rule — what counts, what does not, and when the situation never arose. Seeded from the code's description, never from its repair recommendation, since what to do about a failure is not what a judge decides.
Every judgement is occurred, did not occur, or not applicable.
Folding "the situation never came up" into "it did not happen" makes every rate read low, which is the direction that looks like good news. Deleting the third verdict breaks the import, not a config change.
| Verdict | Counted as | Why |
|---|---|---|
| Occurred | Numerator and denominator | The failure was there. This is the thing being measured. |
| Did not occur | Denominator only | The judge looked and the failure was absent. That is evidence. |
| Not applicable | Neither | The situation never arose, so the judge learned nothing about this code. |
| Skipped or unrecognised | Neither | A trace the model did not rule on is never a negative. Skipped traces and unrecognised labels are each counted on their own. A response that could not be parsed at all lands in skipped, where it is not distinguishable from a trace the model silently left out. |
A judge is a model ruling on one narrow question. The only reason to believe it rules the way a person would is that it was checked against a person who already did it by hand, on the same traces.
An annotation on the trace, tagged to this code. The turn it pointed at is kept, so the page can link back to it.
The trace carries notes, and none of them are for this code. You looked and did not name it.
A trace with no annotations is dropped from the comparison entirely. A human who never opened it did not testify that the failure was absent.
| Number | What it says | How much to trust it |
|---|---|---|
| Recall | Of the traces you flagged, the share the judge also caught. | The headline Your positives are the trustworthy half, so this is the number the judge is judged on. |
| Specificity | Of the traces you cleared, the share the judge also cleared. | Optimistic People tag what they notice. A miss on your side reads as agreement and inflates this. |
| Agreement | Raw share of identical rulings. | Context only Failure modes are long-tailed, so a judge that always says no scores well here. |
| Kappa | Agreement corrected for chance. | Shown as a dash when undefined Two raters who each gave one answer throughout cannot be measured, which is not the same as chance-level agreement. |
source_id for a verdict to be joined to.A validated judge runs across every trace on the evaluation and reports how often the failure was present. One run is one judge against one evaluation.
The denominator is the traces the judge actually ruled yes or no on. Everything else is excluded, and each exclusion is reported as its own count beside the rate.
Every verdict is written in one batch when the run finishes. A run that dies partway leaves the individual verdicts unwritten, and only the run itself is marked failed, so a partial run is not inspectable.
If every trace came back not applicable, the run is stored with no rate at all rather than 0%.
Zero would say "we looked and it never happened". Nobody looked at anything the code applies to, and those are different claims.
This is what the whole loop was for. The same judge runs against a run taken after the repair, and the two rates are put side by side.
A baseline run is the evaluation a repair was launched from. A post-repair run is one carrying the instruction version that repair produced.
Anything else is labelled unrelated and never paired. A delta between two runs that are not comparable is worse than no delta at all.
The panel states the movement in points and names the two instruction versions it was measured across.
If the two evaluations ran different pipeline phases, it says so, because then their traces differ for reasons other than the repair and the movement is suggestive rather than measured.
The loop is honest about its own limits, and they are worth knowing before the first number is quoted.
One person codes, deliberately. Design by committee dilutes the coding and produces categories nobody meant.
And people tag what they notice, so a failure an annotator missed looks like a failure that was absent. That is why recall leads and specificity is labelled optimistic.
Validation says a judge agrees with you on the traces you coded. It does not say the decision rule is a good definition of the failure in general.
That question is what validation exists to keep asking. A judge whose recall drops is telling you the rule has drifted from what you meant.
A rate falling after a repair is consistent with the repair working. It is not proof, especially if the two runs differed in phases, models or traces.
The panel names both instruction versions and warns on a phase mismatch for exactly this reason.
A run produces dozens of conversations and reading them is the price of the method. It is the only step that cannot be delegated to a model without losing the thing it produces.