AESOP evaluation mode · September 2026

Sounds Like AI

Ask an agent ordinary questions. Keep the answers. Score how much they read like a machine wrote them.

Nothing is attacked here and nothing is scored for accuracy. An agent can write badly and still be right.

01 What this measures the short version

Most evaluation modes put an agent under pressure. This one does not.

The agent gets twelve everyday requests drawn from what it was built to do. Explain this. Compare those two. Draft a note. Troubleshoot a setting. No adversary, no conductor, no follow-up designed to trip it. Just the questions a real user would send.

Every reply is then read against a catalogue of 139 writing patterns and given a score from 0 to 100. High means the writing reads human. Low means it reads machine-written.

That is the whole mode. Three things it deliberately is not.

Not an attack

Ordinary questions only

Prompts come from the agent's own stated purpose, varied in shape: explanations, recommendations, troubleshooting, comparisons, drafts.

Nothing tries to make it fail. Red Team does that. This mode is the quiet one.

Not an accuracy score

A correct answer can read badly

Nothing checks whether the reply was right. A correct answer that opens with “Great question!” still loses points for opening with “Great question!”

Right content, wrong register is the thing being found.

Not a verdict

The score describes the writing

It is one signal about one behaviour, reported on its own.

It never feeds the composite score or certification. A safe agent that says “delve” does not get vetoed for it.

02 The run write, ask, scan, adjudicate, score

Five stages, in order. Two of them involve a model. Three are Python.

1. Write the prompts
Twelve ordinary questions drawn from the agent's stated purpose. Not attacks, not edge cases. Varied in shape: explain, recommend, troubleshoot, compare, draft.
An AI writes it
↓
2. Ask the agent
Each prompt goes to the agent once and the reply is kept. No adversary, no conductor, no pressure. One turn, then the transcript is stored.
The agent under test
↓
3. Scan
Every reply is checked against the catalogue in plain Python. Code blocks, inline code, URLs and markdown table rows are blanked first, so a code agent is not failed for its own flags and pipes.
Python
↓
4. Adjudicate
Only matches whose pattern has a stated legitimate reading are shown to a model, which rules each one real or not. It returns verdicts only. Every number stays Python's.
The adjudicator
↓
5. Score
Confirmed matches set the score. Repeating one pattern counts for less than it looks, because the twentieth em-dash says no more than the fourth.
Python
An AI writes or rules The agent under test Python does the counting The adjudicator model
The line to take away. A model is allowed to judge a word in a sentence. It is never allowed to produce a number. The scan is written to the database before the adjudicator is called, so a timeout costs the refinement, not the run.
03 The catalogue 139 patterns, six families

One JSON file is the single source of truth. The detector reads it, the adjudicator's instructions are rendered from it, and the run records its content hash. Change what gets detected by editing the file, not the code.

Every entry also has to answer one question: when is this legitimate? A pattern that cannot answer it is not ready to be in the catalogue.

FamilyPatternsNeeds a rulingWhat it covers
Word choice6047delve, tapestry, leverage, robust, landscape
Sentence constructions266“not X but Y”, “let's dive in”, signposting
Hedging152“it's worth noting”, “could be argued”, adverb crutches
Formatting and rhythm1212em-dash density, tricolon density, uniform paragraphs
Chatbot pleasantries143“Great question!”, “I hope this helps”, “want me to also”
Simulated empathy1212“I hear you”, “sit with that”, “how does that land?”
Total1398257 have no defensible use and cost no model call
Two switches worth knowing about
  • The empathy family is off by default, per agent. “I hear you” is correct in a support agent and a tell in a code assistant. Turn it on where the register is right.
  • The catalogue is English only. That is not a limitation to work around. It is the limit of what this mode can say, and it says so in the report.
04 A match is a question, not a verdict the legitimate reading

A pattern with a legitimate reading is never trusted on a match.

Every entry has to state the reading that makes it legitimate. “Harness” is a tell. “Test harness” is not. A regex cannot tell you which one it found. A person can, and so can a model given the sentence.

So the catalogue splits in two. 82 entries carry a legitimate reading and a match on one goes to the adjudicator, which rules real or not. The other 57 have no defensible use and are trusted on the pattern alone. Those cost no model call at all.

That split is the difference between a confident false positive and a question. “Deliberate” firing on “the board deliberated” is noise. “Delve” is not, and it does not need to be asked about.

What the adjudicator is shown. Not the bare word. The match plus about 140 characters either side, on one line, so the sentence is legible. The verdict has to name the span it is ruling on. A verdict for a span the model was not shown is discarded.
What it can change. Only whether an occurrence counts. It never returns a score, and one that somehow does is thrown away. If it times out, the unadjudicated scan stands and the run still finishes.
05 Density, not presence one em-dash is punctuation

One em-dash is punctuation. Twelve is a tell.

So the detector does not count the first occurrence of a detail pattern. It subtracts an allowance first, based on how long the text is, and only the excess counts. A long document is not punished for using its own punctuation normally.

Then a ceiling: no single pattern can contribute more than three occurrences, no matter how many times it fires. So one punctuation habit cannot outweigh six different real tells sitting next to it. The twentieth em-dash says no more than the fourth, and it is scored that way.

The score has two parts. A presence term, which is how many different patterns are in play. And a density term, which is how thick they run. Density saturates rather than clipping, so a short reply does not hit a cliff at the cap, which is exactly where the cap would bite hardest.

Worked, in order
  • Blank the code. Fenced blocks, inline code, URLs and table rows go before anything is counted
  • Subtract the allowance. What is left over is the count for that pattern
  • Cap at three. One pattern, one ceiling, whatever it does
  • Weight and blend. Presence plus saturating density, subtracted from 100
06 A confident zero is worse than no answer nothing to score

A reply written in another language matches nothing in an English catalogue. Scored naively, it gets a perfect 100 for the worst possible reason.

Runs that cannot be scanned return no score. The report says why. A run that produced no usable replies at all fails outright rather than passing quietly.

Short runs are still scored, but flagged as a thin sample, because twelve one-line answers do not add up to a reading of anything.

And some things cannot be measured in code at all. An agent restating your question as its opening line is a tell, and no counter will catch it. The report names those patterns explicitly, so a clean score is never read as covering something nobody checked.

The principle. A missing answer and a good answer are different things. This mode would rather say no score than say 100.
07 How you read the score the bands

0 to 100, higher is better. Four bands, and the words are the ones the run page uses.

ScoreReads asWhat it means
90 – 100Reads humanThe replies carry almost none of the catalogue. Nothing to look at.
80 – 89Mostly humanA few patterns, most of them ruled legitimate. Worth a skim of the flagged list.
65 – 79Noticeably patternedEnough patterns firing at enough density that a reader would feel it, even without naming it.
Below 65Reads machine-writtenThe register is the model's default, not the agent's purpose.

Every flagged match is shown in the sentence it came from, so you can judge it yourself rather than take the number's word for it. Confirmed matches and ruled-out ones are both listed.

08 What this does not tell you read this part

The score describes the writing, not the answer. It says nothing about whether the reply was correct, useful, or safe.

It is a stylistic signal, not proof that text was machine-written. People reach for stock phrases too, and a human under deadline writes in cliches.

It is advisory. The mode reports its own 0 to 100 score and nothing else. It does not feed the composite score and it does not touch certification.

The limits, stated plainly
  • English only. Other languages match nothing and return no score rather than a false pass
  • Empathy is off by default. Correct in a support agent, a tell in a code assistant. Per agent, your call
  • Some tells are unmeasurable. Restating your question, for one. The report names them so a clean score is not mistaken for full coverage
  • Twelve prompts is a sample. Enough to see a register repeat, not enough to characterise an agent
© 2025–2026 Charlie Fuller All AESOP explainers →