Think about how most AI runs in production today. A support ticket comes in, and we pay a frontier LLM to read it, write a paragraph about it, and finally say “refund”. Then our code parses that paragraph back into a single value for a switch statement.
That’s like hiring a novelist to tick checkboxes.
A large share of the LLM calls I see in client systems work this way. The output is prose, but the result is one of a handful of known answers: which team, which intent, safe or unsafe, allowed or blocked. We’re paying for generation when what we needed was a decision.
A Model That Doesn’t Write
In September, TypeSafe AI released Jev, which it describes as the first “System One” model. The name borrows from Daniel Kahneman’s idea of System One thinking: fast, intuitive judgments, as opposed to slow, deliberate reasoning.
Jev doesn’t generate text at all. You give it some context and a set of typed questions, such as “Which of these five intents is this?” or “Does this need escalation: yes or no?”. It returns typed answers with a confidence score for each one. There’s no paragraph to parse, and no risk of the model wandering off-format.
What the Benchmarks Show
OpenRouter published its own head-to-head tests of Jev 1.13 against Claude Opus 5. The first task was triaging 60 support tickets into five intents, plus an escalation flag:
| Claude Opus 5 | Jev 1.13 | |
|---|---|---|
| Intent correct | 60 / 60 | 59 / 60 |
| Escalation correct | 59 / 60 | 60 / 60 |
| Median latency | ~2.0 s | ~0.19 s |
| Cost per 1,000 tickets | $2.88 | $0.025 |
Accuracy was effectively a tie. OpenRouter’s own summary was “accuracy is a wash”. But Jev was about 10x faster and roughly 115x cheaper.
On a second task, screening 40 messages for prompt-injection attempts, both models scored 40/40. Jev did it at about 1/100th of Opus’s cost.
At 40,000 tickets a month, the triage numbers work out to about $115 a month with Opus versus about $1 with Jev.
The pricing explains why. Jev charges $0.042 per million input tokens, and output is free, because there are no output tokens. TypeSafe describes all the questions in a call as being answered in a single pass, not generated one token at a time.
The Fuller Picture
Small samples flatter everyone, so it’s worth looking at OpenRouter’s larger test as well. It used 3,080 banking queries across 77 intents (the public Banking77 dataset):
- Accuracy: Jev 81.0%, Opus 84.4%. That’s a real 3.3-point gap.
- Speed: Jev was about 13x faster at the median (175 ms vs 2,266 ms).
- Cost: Jev cost about 1/22 of Opus, even with Opus using prompt caching.
The most interesting result is the cascade. When OpenRouter let Jev answer everything it was at least 90% confident about, and sent the rest to Opus, accuracy came back to 84.0% at $0.69 per 1,000 requests. Jev handled 76% of the traffic on its own.
That is the real lesson. The goal isn’t to replace the LLM with something cheaper. It’s to stop sending it the work that doesn’t need it.
Where Decision Models Fit Today
OpenRouter’s guide puts the rule of thumb neatly: “Anything that ends in a switch statement is a Jev question.”
In practice, that covers a lot of production AI:
- Intent routing and ticket triage
- Classification and tagging at scale
- Prompt-injection and safety gates
- Deciding whether an agent may call a tool
- Checking whether an LLM’s answer is supported by your policy text
- Ranking, filtering and lead scoring
Where They Don’t
It’s just as important to know the limits. Based on OpenRouter’s testing, Jev:
- Can’t write anything: replies, documents, summaries and plans still need an LLM.
- Takes text only: no images, audio or PDFs.
- Struggles with arithmetic, counting and date comparisons.
- Loses accuracy when the input contains a lot of irrelevant detail, so trim the context you send.
- Reads criteria literally, so your question and answer options need to be precise.
One more caution. Treat Jev’s confidence score as a ranking signal, not a guaranteed probability. OpenRouter’s two write-ups don’t fully agree on how well-calibrated it is. Set your thresholds using a sample of your own labelled data, not a default.
The Architecture We Recommend
This is the pattern we’re now recommending to clients at Pageup:
graph LR
A["Incoming request"] --> B["Jev decides"]
B --> C{"Confidence above threshold?"}
C -- Yes --> D["Code applies the decision"]
C -- No --> E["LLM or human reviews"]
D --> F["LLM writes text, only where a person will read it"]
F --> G["Jev checks the output"]- Jev decides. Every routing, classification and gating step goes to the decision model first.
- Code applies thresholds. Confident answers are acted on directly. Uncertain ones go to an LLM or a person.
- The LLM writes only where a human will read prose. Customer replies, summaries, documents.
- Jev checks the output. For example, whether a drafted reply stays within your policy text.
The LLM doesn’t go away. It goes back to the job it’s actually good at, which is generating language, and stops billing you for decisions.
The Next Phase of AI Engineering
In my view, 2025 and 2026 were about making models bigger. The next phase is about using the right kind of model for each step. Decisions and generation are different jobs, with different costs and different failure modes. Our architectures should treat them that way.
None of this needs a big rewrite. Start by looking at your logs. Count how many of your LLM calls end with your code picking one value out of a fixed list. That number is usually higher than teams expect, and every one of those calls is a candidate.
Conclusion
If your AI bill is growing faster than your product, the problem may not be the price of the model. It may be the kind of model you’re using for each step.
Split decisions from generation. Let a fast, cheap decision model handle the switch statements, keep the LLM for the words, and measure both on your own data.
If you’d like a second pair of eyes on where your AI spend is going, talk to our team. We’re happy to help you find the calls that should never have been LLM calls.




