Newsroom — Engineering
Jev: the System One model
TypeSafe AI's Jev doesn't generate text. It answers typed questions with calibrated probabilities — fast, cheap judgment for agent stacks. Why the missing layer in AI products isn't a bigger model, it's a faster decider.
By VewMet — October 2026 — 4 min read
Contents
A model that doesn't talk
On 15 September 2026, TypeSafe AI released a model that does something heretical: it doesn't generate text. No chat, no completions, no free-form tool calling. You POST it a state — text or JSON — plus a set of typed questions, and it returns typed answers: a Choice picked from options, a Score on ordered levels, a Noul — a yes/no probability — each with probabilities and a confidence attached.
The company was founded by Diogo Almeida, ex-OpenAI and a co-author of InstructGPT. The product's name carries its thesis twice over. "Jev" nods to Daniel Kahneman's System 1 — the fast, intuitive mode of thinking, set against the slow, deliberate System 2. It also honors William Stanley Jevons, of the Jevons paradox: make decisions cheaper, and the world makes more of them.
Decisions, not prose
Here is the bet: most of what agent stacks actually need from a model isn't prose. It's judgment. Should this ticket be escalated or handled? Which tool should this step call? Which of these five routes should the request take? Which citation actually supports the claim?
Today, teams answer those questions by prompting a giant language model and squinting at the output — parsing a "yes" out of a paragraph, or worse, out of a reasoning trace. It works, expensively and slowly, and the probabilities it returns are ones you can't do arithmetic on.
Jev collapses that into an API shaped like the question. Choice, Score, Noul. Seventy to five hundred milliseconds. About $0.042 per million input tokens, output free. No fine-tuning step — you describe the decision, send the state, get the number back.
What deserves appreciation here is calibration: a model that knows what it doesn't know, and says so in the only language software understands — a number.
One request, three answers
Concretely, a Jev call looks like this. One block of state, three typed questions, each evaluated in parallel against the same state:
POST https://api.typesafe.ai/v1/systemone
Authorization: Bearer <redacted>
Content-Type: application/json
{
"model": "jev-latest",
"state": "Ticket #4821: 'The invoice page throws a 500 when I export to PDF. Third time this month.'",
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should own this ticket?",
"criteria": {
"billing": "Invoices, charges, payments",
"platform": "Errors, crashes, broken exports",
"support": "How-to questions, account help"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Needs attention this week", "Blocking the customer now"]
},
"churn_risk": {
"type": "noul",
"instructions": "This customer is at risk of churning."
}
}
}And back come typed answers — a pick, a position, a probability — each with numbers your code can actually do arithmetic on:
{
"answers": {
"route": {
"choice": "platform",
"probabilities": { "billing": 0.04, "platform": 0.91, "support": 0.05 },
"confidence": 0.93
},
"urgency": {
"score": 2.6,
"legend": ["Routine", "Needs attention this week", "Blocking the customer now"],
"probabilities": [0.08, 0.34, 0.58],
"confidence": 0.71
},
"churn_risk": { "noul": 0.82 }
},
"model": "jev-1.13.0",
"latency_ms": 214
}Illustrative request and response, shaped like TypeSafe's documented API. No parsing a "yes" out of a paragraph — the answer space was declared before the call, so the model cannot return a value outside the schema.
Trained to decide, not to please
The interesting technical claim is in the training. Jev is trained with RLCD — reinforcement learning for calibrated decisions — and explicitly not with RLHF. TypeSafe's thesis is blunt: RLHF, the technique that made chatbots polite and helpful, makes models conservative and miscalibrated for programmatic use. A model trained to please a human rater learns to hedge, to flatter, to sound sure. Software doesn't need to be charmed. It needs a probability it can trust.
Whether RLCD is the answer, nobody outside the company can verify yet. But the diagnosis rings true to anyone who has tried to extract a reliable yes-or-no from a chat model at 2 a.m. before a launch. Pleasant is not the same as correct, and confident is not the same as calibrated.
The numbers, with attribution
TypeSafe's reported benchmarks — their numbers, not independently verified: 81.0% on Banking77 against Claude Opus 5's 84.4%. Within striking distance, at roughly 1/22nd the cost and about 13x the median speed. The company positions it for the unglamorous load-bearing jobs: routing, tool-call gating, reranking, citation checks, injection screening.
Read that list again. Those are decisions every agent stack makes dozens of times per run, usually with a sledgehammer — one giant model invoked for judgments a tenth its size could make.
And a word of honesty, which this industry could use more of: the frontier claims — calibration quality, intelligence-per-dollar — are company assertions. Treat them as a direction, not a fact. We'll believe the independent evals.
The missing layer
Here is the builder's view from VewMet. The last two years gave us System 2 in a box: large models that reason slowly and deliberately, and they are remarkable at it. But an agent stack built only of slow thinkers has a hole in the middle. Every step of a run is a small decision — act or escalate, this tool or that one, this route or the other — and making each one with a deliberative giant is like hiring a philosopher to direct traffic.
What has been missing is System 1 for software: fast, cheap, calibrated judgment, available as infrastructure. That is the layer Jev is trying to be. And the Jevons logic cuts in its favor: when each decision costs a fraction of a cent and returns in a tenth of a second, you stop rationing decisions. You instrument everything. Routing, gating, screening — the connective tissue of agents — gets a dedicated organ instead of borrowed time on the big brain.
Nobody should be reading reasoning traces to decide whether to escalate a support ticket.
Where it lands for us
This is close to how we already think at inkk.ai. A collaborative writing workspace makes thousands of small routing judgments: which provider's model should handle this request, which marketplace agent fits this task, what gets suggested and what stays quiet. Multi-provider model routing is already in the product. The open question is always which judgment is worth spending where — and a calibrated decider changes that math.
The trend worth appreciating is cheap, trustworthy machine judgment moving into the plumbing — not as a demo, but as a line item cheaper than the logging around it. If Jev's calibration claims hold up under independent testing, the agent stack gets a new division of labor: the big models think, the small ones decide, and the human stays in charge of what matters.
We'll be watching the evals, not the launch thread.
Keep reading
AI — October 2026
AI for computer use
AI agents can now operate a computer the way a person does — seeing the screen, moving the mouse, typing. What that changes, where it breaks, and why the agent belongs inside the work, not on top of the operating system.
Product — February 2025
inkk.ai Marketplace is live
Our AI agent marketplace launched in February 2025 — AI agents that take care of all your writing needs.
Product — October 2026
OpenAI just announced Pages. We've been shipping it since February.
OpenAI's Pages is a collaborative AI document editor. Ours has been live since February — with caret-level attribution, in-document calls, and zero vendor lock-in.
Product
VewMet Calls
Imagine AI joining your calls as a participant — listening, moderating, and contributing in near real-time. Launching soon.