Jev is now available in LangSmith Evals

Jev is now available as a judge for evaluations in LangSmith, according to an announcement dated September 21, 2026 by Winston Huynh. TypeSafe is now a model provider in LangSmith, with Jev available as a model under the name jev-latest.
LangChain describes Jev as a System One model, a class of AI models that evaluates a state and returns typed answers and probabilities rather than generating text. For evals, the state can be an agent trace, a single message, or other context. Questions define the criteria, and Jev can answer three types: a noul returns a yes/no probability, a choice picks one option from a set, and a score rates the state on an ordered scale. Each answer comes back typed instead of as generated text converted into structured output.
According to TypeSafe AI, Jev is up to roughly 450x cheaper and 200x faster than comparable LLMs on classification tasks, and it can evaluate multiple questions about the same state in parallel. LangChain states that LLM judges remain preferable for open-ended criteria where written reasoning alongside a verdict is wanted, and notes fine-tuned and open models can also serve as lower-cost judges.
LangChain said it tested Jev against GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6. Jev matched a human reviewer on every decision, with 92-913x lower variance than the LLM judges, and averaged 0.44 seconds per call versus 2.16-2.83 seconds. At $0.00035 per call, the full set of judgments cost $0.34 with Jev, against $0.39 with GPT-5.6 Luna, $2.90 with GPT-5.6 Terra, and $28.17 with Claude Sonnet 4.6. LangChain described this as one test on one agent.
Setup follows the same path as an LLM-as-a-judge evaluator: add a TypeSafe API key in Provider secrets, add an evaluator from the Evaluators tab, select TypeSafe as provider and jev-latest as model, define the state by mapping run or thread variables, then add one question per criterion. Jev-as-a-judge is available in LangSmith today.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests.