Ask an agent ordinary questions. Keep the answers. Score how much they read like a machine wrote them.
Nothing is attacked here and nothing is scored for accuracy. An agent can write badly and still be right.
Most evaluation modes put an agent under pressure. This one does not.
The agent gets twelve everyday requests drawn from what it was built to do. Explain this. Compare those two. Draft a note. Troubleshoot a setting. No adversary, no conductor, no follow-up designed to trip it. Just the questions a real user would send.
Every reply is then read against a catalogue of 139 writing patterns and given a score from 0 to 100. High means the writing reads human. Low means it reads machine-written.
That is the whole mode. Three things it deliberately is not.
Prompts come from the agent's own stated purpose, varied in shape: explanations, recommendations, troubleshooting, comparisons, drafts.
Nothing tries to make it fail. Red Team does that. This mode is the quiet one.
Nothing checks whether the reply was right. A correct answer that opens with “Great question!” still loses points for opening with “Great question!”
Right content, wrong register is the thing being found.
It is one signal about one behaviour, reported on its own.
It never feeds the composite score or certification. A safe agent that says “delve” does not get vetoed for it.
Five stages, in order. Two of them involve a model. Three are Python.
One JSON file is the single source of truth. The detector reads it, the adjudicator's instructions are rendered from it, and the run records its content hash. Change what gets detected by editing the file, not the code.
Every entry also has to answer one question: when is this legitimate? A pattern that cannot answer it is not ready to be in the catalogue.
| Family | Patterns | Needs a ruling | What it covers |
|---|---|---|---|
| Word choice | 60 | 47 | delve, tapestry, leverage, robust, landscape |
| Sentence constructions | 26 | 6 | “not X but Y”, “let's dive in”, signposting |
| Hedging | 15 | 2 | “it's worth noting”, “could be argued”, adverb crutches |
| Formatting and rhythm | 12 | 12 | em-dash density, tricolon density, uniform paragraphs |
| Chatbot pleasantries | 14 | 3 | “Great question!”, “I hope this helps”, “want me to also” |
| Simulated empathy | 12 | 12 | “I hear you”, “sit with that”, “how does that land?” |
| Total | 139 | 82 | 57 have no defensible use and cost no model call |
A pattern with a legitimate reading is never trusted on a match.
Every entry has to state the reading that makes it legitimate. “Harness” is a tell. “Test harness” is not. A regex cannot tell you which one it found. A person can, and so can a model given the sentence.
So the catalogue splits in two. 82 entries carry a legitimate reading and a match on one goes to the adjudicator, which rules real or not. The other 57 have no defensible use and are trusted on the pattern alone. Those cost no model call at all.
That split is the difference between a confident false positive and a question. “Deliberate” firing on “the board deliberated” is noise. “Delve” is not, and it does not need to be asked about.
One em-dash is punctuation. Twelve is a tell.
So the detector does not count the first occurrence of a detail pattern. It subtracts an allowance first, based on how long the text is, and only the excess counts. A long document is not punished for using its own punctuation normally.
Then a ceiling: no single pattern can contribute more than three occurrences, no matter how many times it fires. So one punctuation habit cannot outweigh six different real tells sitting next to it. The twentieth em-dash says no more than the fourth, and it is scored that way.
The score has two parts. A presence term, which is how many different patterns are in play. And a density term, which is how thick they run. Density saturates rather than clipping, so a short reply does not hit a cliff at the cap, which is exactly where the cap would bite hardest.
A reply written in another language matches nothing in an English catalogue. Scored naively, it gets a perfect 100 for the worst possible reason.
Runs that cannot be scanned return no score. The report says why. A run that produced no usable replies at all fails outright rather than passing quietly.
Short runs are still scored, but flagged as a thin sample, because twelve one-line answers do not add up to a reading of anything.
And some things cannot be measured in code at all. An agent restating your question as its opening line is a tell, and no counter will catch it. The report names those patterns explicitly, so a clean score is never read as covering something nobody checked.
0 to 100, higher is better. Four bands, and the words are the ones the run page uses.
| Score | Reads as | What it means |
|---|---|---|
| 90 – 100 | Reads human | The replies carry almost none of the catalogue. Nothing to look at. |
| 80 – 89 | Mostly human | A few patterns, most of them ruled legitimate. Worth a skim of the flagged list. |
| 65 – 79 | Noticeably patterned | Enough patterns firing at enough density that a reader would feel it, even without naming it. |
| Below 65 | Reads machine-written | The register is the model's default, not the agent's purpose. |
Every flagged match is shown in the sentence it came from, so you can judge it yourself rather than take the number's word for it. Confirmed matches and ruled-out ones are both listed.
The score describes the writing, not the answer. It says nothing about whether the reply was correct, useful, or safe.
It is a stylistic signal, not proof that text was machine-written. People reach for stock phrases too, and a human under deadline writes in cliches.
It is advisory. The mode reports its own 0 to 100 score and nothing else. It does not feed the composite score and it does not touch certification.