An AI is good at reading a messy situation and forming a judgement. It is unreliable at arithmetic, at following the same rule twice, and at saying the same thing in the same way.
So AESOP stopped asking it to do those things. The result is a build that is more predictable, wastes less, and costs less to run.
What Jev is. Jev is a model made by TypeSafe. Most AI answers by writing; Jev answers by picking from a list you hand it, and says how confident it is in the pick. AESOP gives it the situation and the options. Everything after that is code.
Every stage of a build has to make decisions. Before Jev, AESOP asked the build model to write an assessment, then paid a second program to read that assessment back and guess what it meant. Two problems: the program sometimes guessed wrong, and you were paying twice for one decision.
Now the AI does the part it is actually good at — reading the situation and picking a judgement from a fixed set of options — and hands every number, rule and threshold to ordinary code. Three things improved.
One decision is now made in one place. Before, two separate copies of the same logic could read one answer and disagree — and that answer was what a person reviewed before approving the work.
And you can see why. Every threshold is a number in code, not a sentence in a prompt.
The system was re-sending the same document to the AI on all seven steps, and paying full price every time. It now sends it once and the AI re-reads it from its own cache.
Roughly a third of a build's input cost was going to repetition.
A score used to be pulled out of a paragraph by pattern-matching. One test document that said “we reviewed 3/100 documents” was recorded as a score of 3.
Code owns the arithmetic now. The AI supplies the judgement; it does not supply the total.
None of these are hypothetical. Each one ran in production and was found by comparing what the code did against what it was supposed to do.
| What went wrong | What it meant |
|---|---|
| Two copies of one rule | The same written assessment was read two different ways by two parts of the system. One said the answer was stop; the other said continue. That answer decides whether a person is asked to approve the work, so the difference mattered. |
| A number found in a sentence | A program scanned documents for any number-over-100. A sentence reading “we reviewed 3/100 documents” was recorded as a score of 3. It looked like a real result. It was a coincidence of wording. |
| The same facts written three times | One structured record was turned into prose, stored, turned into prose again by the next step, and stored again. Three copies of one truth, each free to drift from the others. |
Take the first decision the system makes: is this project fit to build? It scores five things, then compares the total against a pass mark and a lower bar. Separately, a short list of conditions fails it outright no matter how well it scores.
Before, all of that lived in the instructions handed to the AI, which then wrote an answer, and code read the answer back out. Now the AI reports the five scores and names any disqualifying condition it found. Everything after that is arithmetic, and arithmetic belongs in code.
The instructions the AI reads are now generated from the same numbers the code checks. That sounds like a detail. It is the whole thing.
Before, the pass mark appeared once in the AI's instructions and again somewhere in the code. Two facts that happened to agree. Change one and nothing tells you the other is now wrong. Now there is one number, and the instructions are printed from it. They cannot disagree.
This is the part with a number attached, so it is worth explaining properly.
A build runs in seven steps. Each step needs the same background document — the requirements, the interview notes, the reference material. The system was sending that whole document to the AI again on every single step, at full price, even though none of it had changed.
AI providers offer a way out: if the text at the start of a request is byte-for-byte identical to last time, they charge a fraction of the price to re-read it. A tenth, typically. The catch is that everything before the marker has to match exactly — one character out of place and you pay full price again, with no error to tell you.
Something was always out of place. Each step was given a slightly different set of instructions at the very start, which meant the shared document behind it never counted as “the same” and was charged in full every time.