Jev: An AI Model That Returns Types, Not Text

For three years, nearly every new model has been sold on being a better talker. Jev, which came out of stealth on September 15, goes the other way: it won’t chat with you, and it never emits a line of natural language. You hand it a chunk of state plus a structured question, and it returns a typed result carrying a probability distribution and a confidence value — a boolean probability, an option picked from a list you supplied, or a score on a scale you defined. TypeSafe AI, the company behind it, positions Jev as AI for software to consume, not AI for people to read.
Why this company gets a look
TypeSafe AI is a San Francisco company founded in 2024. Co-founder and CEO Diogo Almeida spent roughly four years at OpenAI working across RLHF, InstructGPT, ChatGPT and GPT-4; co-founders Erik Gafni and Sasha Sheng bring comparably deep AI backgrounds. After two years in stealth, the company announced both its exit and a $40 million seed round on September 15, led by DCVC. Forbes, citing a person familiar with the deal, reported a post-money valuation around $200 million. DCVC partner James Hardiman framed the bet as TypeSafe taking on “one of the biggest remaining challenges in AI” — making models reliable enough to be embedded inside other software at scale.
Almeida’s assessment of today’s generative models is blunt: “We’ve been optimizing for humans, and we’re superhuman at pleasing humans.” The implication being that a model trained to delight a human reader is not the same artifact as one you want sitting inside a pipeline that needs stable output.
What a “System One model” means
TypeSafe calls Jev the first “System One model,” borrowing Kahneman’s fast/slow split: System 1 handles quick, intuitive judgments that don’t need reasoning spelled out. That’s the slice Jev is claiming — atomic questions a knowledgeable person could answer at a glance, not tasks that require extended deliberation.
The name is its own small joke. Jev comes from William Stanley Jevons, the 19th-century economist behind the Jevons paradox: make a resource cheaper to use and total consumption goes up, not down. TypeSafe is clearly betting that once judgment gets cheap enough, it ends up in far more places than it is today.
Three primitives: Choice, Score, Noul
Jev’s API accepts exactly three kinds of question, which the docs call primitives:
| Primitive | The question | What comes back |
|---|---|---|
| Choice | Pick one option from a list | The selected option, a probability per option, a confidence value |
| Score | Rate the state against a rubric | A score, a probability per level, a confidence value |
| Noul | Is this statement true? | A single probability from 0 to 1 |
Choice supports up to 255 discrete options and uses two-stage independent scoring. Noul has no separate confidence field, because the 0-to-1 value is the certainty — 0.5 means the model genuinely can’t call it. For Choice and Score, confidence is derived from the probability distribution: a concentrated distribution yields high confidence, a flat one yields low confidence.
Every question runs in parallel and in isolation against the same state. That’s a meaningful design constraint: you cannot ask it to “weigh these three factors and give me a verdict.” TypeSafe’s guidance is to decompose multi-factor questions into separate ones and recombine the answers in your own code.
That’s also what confidence is actually for. TypeSafe’s training method, RLCD (Reinforcement Learning for Calibrated Decisions), aims to make the returned confidence track real accuracy rather than producing the overconfidence it argues traditional RLHF bakes in. Once calibration holds, you can threshold in code: above some confidence, proceed automatically; in the middle, gather more information; below it, escalate to a human.
The speed and price claims, and the discount to apply
TypeSafe’s numbers are genuinely eye-catching. End-to-end latency lands between 70 and 500 milliseconds, against 3 to 329 seconds for frontier conversational models on comparable work. Pricing is $0.042 per million input tokens with output unmetered — the model uses hardware-aware parallel sampling rather than autoregressive token generation. Translated into actual workloads, SiliconANGLE’s comparison is $0.39 per 1,000 workflows for Jev, versus $3.31 on OpenAI’s GPT-5.6 Luna and $19.49 on Anthropic’s Claude Haiku 4.5. TypeSafe’s headline multiples are up to 193.6× faster and 444.6× cheaper.
Several discounts apply to all of that, and TypeSafe names most of them itself. The benchmark suite was built by TypeSafe’s own capabilities team, and the “reference answer” is the average of GPT-6 Astra and Fable 5.1. The company states outright that it expects these gains “to sit at the high end of real use” — meaning your own integration probably won’t look this good. More candidly still, TypeSafe says it cannot prove the current price is unsubsidized. As of now, no independent third party has reproduced the figures.
What “zero hallucinations” does and doesn’t mean
“Zero hallucinations” appears in TypeSafe’s marketing, but the claim is narrower than it sounds. Because output is structurally constrained to the type you declared, Jev guarantees the return value matches your schema — nothing unparseable, no bolted-on JSON repair or guardrail wrapper needed. But TypeSafe spells out the corollary too: the answer itself can still be wrong. Schema-correct is not judgment-correct, and the company is careful to keep those two separate.
How you’d actually integrate it, and the current limits
Jev is hosted-API-only today, behind an early-access waitlist, with SDKs for Python (3.10+) and JavaScript. Model weights and parameter counts are unpublished, and there’s no self-hosting option. Two distribution channels already exist: Cloudflare Workers AI serves it under the model ID typesafe/jev with a documented 32K-token context window, and on the Java side, Spring AI published an integration write-up on September 21. The current stable release is jev-1.13.0.
Architecturally, Jev is transformer-based but reportedly trained exclusively on synthetic data. The technical details are unpublished, and some observers suspect an open-weight LLM underneath — something TypeSafe hasn’t confirmed.
When to reach for it, and when not to
The use cases TypeSafe and its coverage list all share one shape: judgments that happen constantly, are individually simple, and turn brittle when you hand-code the rules. Support ticket classification and routing, lead scoring, trust-and-safety review, document classification, candidate filtering, invoice review, security alert triage — even fire-risk assessment in insurance underwriting. One increasingly common pattern is using it as a verification layer over another LLM or AI agent’s output, since that review work is itself high-frequency, atomic, and only needs a boolean or a score.
Conversely, anything that needs extended reasoning, a written justification, generated content, or a single verdict weighing several independent factors at once is outside Jev’s lane — the docs tell you to decompose the question or hand it to a conversational model instead.
For developers, what Jev is really coming for isn’t GPT. It’s the pile of glue code sitting between hard-coded if-else rules and “call an LLM and then try to dig the JSON back out.” That’s a real pain worth a better answer. But until independent evaluations show up, those 400× figures are best treated as a vendor’s best case rather than the basis for your budget.



