AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Can Jev Be a Better Agent Evaluator?

Collected Oct 1, 2026

LangChain published an experiment comparing TypeSafe AI's Jev, a "System One" model, against LLM judges for agent evaluation. The comparison covered accuracy, repeatability, latency, and cost, using GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as baselines. LangChain framed the work as testing whether System One models could form a third kind of agent evaluator alongside code-based evaluation and LLM-as-a-judge.

According to LangChain, Jev does not generate text; it evaluates typed questions against structured state and returns typed answers with probabilities. TypeSafe AI says System One models are built to make fast, structured decisions software can use directly, and claims up to 200x faster inference and 400x lower cost than comparable LLMs on classification tasks. Jev supports choice, score, and Noul (yes/no probability) question types, and multiple atomic questions can be evaluated in parallel against the same state.

The target agent was built with LangChain's open source Deep Agents harness, and the test set was defined as a LangSmith dataset of five weather requests. Each judge evaluated five captured runs with two signals: quality, a continuous score, and does_pass, a binary decision. A human reviewer labeled each fixed response against the same rubric as an oracle. LangChain calculated per-case variance across 100 repetitions and agreement with the oracle.

For the binary does_pass score, LangChain reported Jev matched the oracle on all 500 repeated decisions. Terra matched on 99.8%, Luna on 96.4%, and Claude on 80.0%. On quality-score variance, Jev had the lowest observed mean per-case variance at 0.0000149; LangChain reported Luna 433x higher, Terra 913x higher, and Claude 92x higher. Jev averaged 0.44s and $0.00035 per call, with a total of $0.34 versus $28.17 for Claude, according to the report.

LangChain stated the experiment cannot explain why Jev's scores varied less, calling one hypothesis that the models are optimized for different output types, and describing the result as observational rather than evidence that Jev's training objective caused the lower variance. It said results are promising but early and still need to carry over to other agents and production workflows, noting low cost can amplify mistakes and engineers still need human review and judge alignment. The repository, model access details, and package versions were listed for reproducibility.

Read at LangChain

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

We tested using Jev-as-a-Judge against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation.