AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Jev introduces a new shape of LLM - System One, aka Decision Models

Collected Sep 30, 2026

TypeSafe AI last week unveiled Jev, which the company describes as its first example of a new category of model it calls "System One models". Simon Willison, writing about the release, said he prefers the term "decision models", a name he attributes to Maggie Appleton.

Jev accepts text inputs but instead of returning text, it returns floating point numbers corresponding to categories, yes/no questions, ratings and associated confidence scores. TypeSafe describes it as "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out".

The model is positioned as fast and cheap. Regular LLMs are priced on input and output tokens, with output generally charged at higher rates; Jev charges only for input, with output free, and its first model is priced at $0.042 per million tokens. Willison notes this is cheaper than OpenAI's GPT-5 Nano at $0.05 per million.

Users compose a "state" object containing a string, array of strings, or name-value pairs, then send it to the API with one or more questions. Three question types are supported: Yes/No questions, which Jev calls "Noul" questions - the CEO confirmed on Hacker News this is short for Bernoulli, from the Bernoulli distribution, returning a float between 0 and 1; choice questions, returning a confidence score plus a probability distribution across provided options; and score questions, where numeric levels with descriptions yield a float along that range. Questions are evaluated in parallel.

Jev 1.13 documentation lists current weaknesses: numbers, dates and "adversarial content". Willison suggests decision models suit classification tasks such as spam detection, label suggestion, prioritization and ranking, and says he has experimented with search reranking, scoring 100 BM25 candidates for relevance.

He also raises concerns that Jev represents a further move toward black box machine learning, since it returns only a floating point number with no justification. He says bias concerns should be front and center, adding that he hopes nobody uses Jev to rank job applicants, and describes an experiment scoring San Francisco Bay Area cities on whether they were a "Good city?", which rated Cupertino top and East Palo Alto bottom. He argues evals and structured experiments matter more than for regular LLM projects, and that Jev's low cost makes large experimental runs inexpensive.

Community uses have appeared in the days since release. These include jevchat by Kyle Pena, which Willison calls a terrible chat model; jev-leftpad by Fatih Kadir Akın; and jev-2048 by Andy Gayton. Open weight recreations include Kev, built on Qwen 3.5 at 0.8B, 4B and 9B sizes, with a JevBench benchmark created to compare "Jev-class decision models". Willison also released llm-typesafe, a plugin adding Jev support to his LLM CLI tool and Python library.

Read at Simon Willison

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Last week TypeSafe AI unveiled Jev , their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is a better name for these). Jev is an interesting variant on the usual LLM format: it still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores. TypeSafe describe Jev like this: Think of Jev as a frontier-intelligence fun