There is a new type of model on the block. And it claims to have solved hallucinations.
TypeSafe AI launched Jev on 15 September 2026. They are referring to it as a System One model, borrowing from psychologist Daniel Kahneman’s parlance for fast, intuitive responses (as opposed to System Two, which concerns itself with slower, deliberate reasoning).
TypeSafe says that the model has been trained on synthetic data using RLCD (Reinforcement Learning for Calibrated Decisions). Caveat: they have not published any papers, so independent verification remains a challenge.
In Why do LLMs love corrective juxtaposition?, I argued that RLHF leaves fingerprints on how models write, because human raters reward certain words and structures and the model learns to reach for them.
Diogo Almeida, the founder of TypeSafe, spent four years at OpenAI working on RLHF, InstructGPT, ChatGPT and GPT-4. He left in 2024. His explanation:
We’ve been optimizing for humans and we’re super human at pleasing humans.
This critique goes a level deeper than my assertion. If raters reward what pleases them, and certainty signals competence, a model trained on approval tends to sound sure of itself even when it is not.
1. A judgement layer for software
Frontier labs are chasing AGI and releasing more capable models that perform complex and fuzzy tasks. But the ballooning token costs of such models make widespread adoption harder. Jev aims to bring higher intelligence capabilities to the long tail of decisions inside deterministic programs.
Jev is an optimised judgement layer for software programs. It does not generate text. Its judgements lie within a set of available actions, and the conditions for taking them, that the developer specifies.
It has three building blocks:
choice
score
noul
Here is how it works.
Consider a customer support request: “my flight got cancelled, can I get a refund?”
Choice answers “what is the main request?” The developer feeds the model with predefined choices in a key-value setup, and the model identifies the intent:
“refund”: “Wants money back.”
“rebooking”: “Wants a replacement flight.”
“information”: “Asking for info only.”
The choice options come back with probabilities. If refund comes back with 0.85, rebooking 0.1 and information with 0.05, application code can use this to route the ticket straight to the refunds queue.
Score answers “which level?” for an underlying condition that would help fulfil the support request.
In this case, the underlying condition could be “how frustrated is the customer?” Plausible levels could be:
“calm and neutral”
“concerned but civil”
“very angry”
As with choice, each level comes back with a probability, which the application program can use to decide severity.
Noul is the model’s probability that the answer to a yes/no question is yes. For the above-mentioned support request, it could be the question “is the customer eligible for a refund?”
Since it is a calibrated probability, it lies between 0 and 1. Calibrated means that of all the times the model says 0.9, it should be right about 9 times in 10. A noul value of 0.9 suggests that the model is fairly certain that the customer is eligible. Whereas a value of 0.5 suggests that the model is not confident enough to be relied on solely and human intervention might be needed.
2. Value accrual
The commercial value comes from making uncertainty actionable at much lower costs and much higher speeds. The use cases being pitched include LLM safety monitoring, model routing, and workflow automation where cheap real-time classification is the constraint. The common feature is repetition. A small saving matters when the same decision happens millions of times.
It works commercially because:
Cost is $0.042 per million input tokens, with output tokens free
Reading the same context once and asking several questions about it is cheaper
TypeSafe reports end-to-end latency of 70 to 500 milliseconds
Whether value accrues to TypeSafe is an open question. Judgement is contextual, so a large part of the leverage stays with organisations and enterprises. Jev is competing against traditional LLMs and classifiers, and against smaller open-source models fine-tuned for an enterprise’s context. Its competitive advantage can only come from a combination of accuracy and cost that is hard for either of these approaches to replicate. That will be hard, because the options Jev chooses from are still defined by the developer orchestrating the system, and the model is only as good as the options it is given. Jev can prevent hallucination but not wrong selections.

