Making the build predictable · September 2026

Jev judges. Python does the math.

An AI is good at reading a messy situation and forming a judgement. It is unreliable at arithmetic, at following the same rule twice, and at saying the same thing in the same way.

So AESOP stopped asking it to do those things. The result is a build that is more predictable, wastes less, and costs less to run.

01 What changed the short version

What Jev is. Jev is a model made by TypeSafe. Most AI answers by writing; Jev answers by picking from a list you hand it, and says how confident it is in the pick. AESOP gives it the situation and the options. Everything after that is code.

Every stage of a build has to make decisions. Before Jev, AESOP asked the build model to write an assessment, then paid a second program to read that assessment back and guess what it meant. Two problems: the program sometimes guessed wrong, and you were paying twice for one decision.

Now the AI does the part it is actually good at — reading the situation and picking a judgement from a fixed set of options — and hands every number, rule and threshold to ordinary code. Three things improved.

Predictable

The same input gives the same answer

One decision is now made in one place. Before, two separate copies of the same logic could read one answer and disagree — and that answer was what a person reviewed before approving the work.

And you can see why. Every threshold is a number in code, not a sentence in a prompt.

Efficient

Less work is thrown away

The system was re-sending the same document to the AI on all seven steps, and paying full price every time. It now sends it once and the AI re-reads it from its own cache.

Roughly a third of a build's input cost was going to repetition.

Auditable

Numbers are counted, not read

A score used to be pulled out of a paragraph by pattern-matching. One test document that said “we reviewed 3/100 documents” was recorded as a score of 3.

Code owns the arithmetic now. The AI supplies the judgement; it does not supply the total.

The line to take away. This was never about removing the AI. It was about being precise about which half of a decision belongs to it. Jev reads the situation. Python does the math. Mix those two up and you pay twice — once for the judgement, and again to guess what it meant.
02 Why the old way was expensive three things that actually went wrong

None of these are hypothetical. Each one ran in production and was found by comparing what the code did against what it was supposed to do.

Before Jev · a judgement, written as an essay A decision has to be made The AI writes prose a written assessment Code guesses pattern-matching on words The decision a person reviews it With Jev · a judgement, stated as a value A decision has to be made Jev picks a label from a fixed list the judgement is still its job Python does the math thresholds, weights, rules the arithmetic is code's job The decision a person reviews it spec field verdict The top row has no schema. Jev can phrase the same judgement three ways and all three are “correct.” The bottom row has one shape, so a missing answer is a signal the code can act on rather than guess past.
A decision point Jev — reading and judging Python — arithmetic and rules What a person reviews
Both rows end at the same approval step. Only one of them can explain why it arrived.

The three failures, in plain terms

What went wrongWhat it meant
Two copies of one rule The same written assessment was read two different ways by two parts of the system. One said the answer was stop; the other said continue. That answer decides whether a person is asked to approve the work, so the difference mattered.
A number found in a sentence A program scanned documents for any number-over-100. A sentence reading “we reviewed 3/100 documents” was recorded as a score of 3. It looked like a real result. It was a coincidence of wording.
The same facts written three times One structured record was turned into prose, stored, turned into prose again by the next step, and stored again. Three copies of one truth, each free to drift from the others.
03 How a decision is made now Jev judges. Python does the math.

Take the first decision the system makes: is this project fit to build? It scores five things, then compares the total against a pass mark and a lower bar. Separately, a short list of conditions fails it outright no matter how well it scores.

Before, all of that lived in the instructions handed to the AI, which then wrote an answer, and code read the answer back out. Now the AI reports the five scores and names any disqualifying condition it found. Everything after that is arithmetic, and arithmetic belongs in code.

Jev reads the brief scores five things and flags anything that disqualifies it The scoring code owns the pass mark, the weights and the order they apply in, and they are checked every time it starts FAIL a disqualifying condition was found the score is still shown alongside it NOT YET something is missing, or nobody has been named as responsible PASS or FAIL the weighted total against the pass mark and the lower bar disqualified missing input weighted total A condition the code does not recognise is ignored rather than trusted, so a made-up reason cannot fail a build. A pass on incomplete evidence is the failure that matters, so a missing answer never produces a confident yes.

Why this is more predictable

The instructions the AI reads are now generated from the same numbers the code checks. That sounds like a detail. It is the whole thing.

Before, the pass mark appeared once in the AI's instructions and again somewhere in the code. Two facts that happened to agree. Change one and nothing tells you the other is now wrong. Now there is one number, and the instructions are printed from it. They cannot disagree.

And a test that holds it there
  • The duplicate is gone and cannot come back. Both parts of the system now use one function, and a test asserts they use the same one rather than that they happen to agree on some examples
  • The instructions are checked too. A test fails if the old prose description ever reappears, so the two cannot drift apart again quietly
  • Roughly 100 separate checks now cover this area, most of them written to fail if the old behaviour returns
04 What it saved paying once instead of seven times

This is the part with a number attached, so it is worth explaining properly.

A build runs in seven steps. Each step needs the same background document — the requirements, the interview notes, the reference material. The system was sending that whole document to the AI again on every single step, at full price, even though none of it had changed.

AI providers offer a way out: if the text at the start of a request is byte-for-byte identical to last time, they charge a fraction of the price to re-read it. A tenth, typically. The catch is that everything before the marker has to match exactly — one character out of place and you pay full price again, with no error to tell you.

Something was always out of place. Each step was given a slightly different set of instructions at the very start, which meant the shared document behind it never counted as “the same” and was charged in full every time.

Before · the start of the request changed, so nothing was ever cheap the repeatable part Step instructions a little different every step The shared document identical every time — but never counted as such What this step adds changes every step, and always will The marker sat after the shared document, so the part it covers already included the bit that changed. Every step paid full price. Measured on a real build: step 1 was charged full price for 14,208 words it had already been sent. The next step got only 256 words at the cheap rate. After · the start never changes, so the rest is cheap the repeatable part A fixed block, built once identical on every step, by construction the system never rebuilds it per step The document and the interview notes What this step adds sits outside the repeatable part Steps 1 through 6 re-read all of it at a tenth of the price measured on a real build: step 0 sent 15,795 words; step 1 re-read all 15,795 cheaply Step 0 sends it at full price. Before the measurement existed, that was an assumption; it is an observation now.
The number. Roughly a third of a build's input cost was being spent re-sending text the provider had already seen. On a fullstack build that is about 40,000 words per run that no longer need to be paid for at full price. It is smaller than it sounds — this is not where most of the money goes — but it is real, and it is now measured rather than assumed.
The part that looks like a mistake. To make this work, the system now sends more text on some steps than it used to. There is a minimum amount a provider will bother caching; below it, the cheap rate silently does not apply and nothing tells you. So the fixed block is deliberately padded past that line. Steps pay a little more up front, and the whole thing behind it becomes cheap. It also changes what the AI reads, which has not yet been verified on a real build. The code says so, next to the decision.
© 2025–2026 Charlie Fuller All AESOP explainers →