Ollama now supports Jev-style decision models
Ollama now supports decision models based on TypeSafe's Jev API, according to the company. The feature is available as of Ollama 0.35 through a new /v1/systemone endpoint.
Decision models take text as state along with a set of named questions and answer them all in one request, providing yes-or-no answers, choices and scores with a probability for every option. Ollama listed ticket triage, model routing, and content or safety moderation as tasks suited to fast decisions.
Three decision models are available today. Nimble is an open-source 9B parameter decision model developed by Bespoke Labs. Tev1 is an experimental 4B decision model from Together AI, and tev1:0.8b is an experimental 0.8B decision model from Together AI. Ollama said more decision models are coming soon, including models served by its cloud.
Ollama reported that decision models run at no additional cost and with lower latency when run locally, because requests do not travel over a network. It said Nimble 9B averaged 91ms per decision in a Pac-Man example when running locally on an M5 Max.
Ollama said Nimble and Tev1 were evaluated on mean accuracy across 13 public data sets with human labels, covering 3,880 decisions, and that Jev 1.13 is from Bespoke Labs' published run on the same decisions. It pointed to Ollama evaluation results and a benchmark suite.
Users can download or upgrade Ollama and pull a decision model, such as by running ollama pull nimble. Requests can be made via curl or through TypeSafe's official Python SDK.
Ollama described this as the first of many releases adding decision model support. Future updates will include faster performance on Apple Silicon powered by MLX and more models specializing in different kinds of decision making.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Ollama now supports decision models, based on TypeSafe's Jev API for fast, typed decisions. Decision models can now be run at no cost with low latency. Based on text, decision models answer yes-or-no questions, choices and scores about it, with a probability for every option.