OpenAI gpt-oss-safeguard
Ollama announced it is partnering with OpenAI and ROOST (Robust Open Online Safety Tools) to bring the latest gpt-oss-safeguard reasoning models to users for safety classification tasks. The models come in two sizes, 20B and 120B, and are permissively licensed under Apache 2.0.
Users can run the models through Ollama with the commands "ollama run gpt-oss-safeguard:20b" and "ollama run gpt-oss-safeguard:120b".
According to the announcement, the models are trained and tuned for safety reasoning, intended for use cases including LLM input-output filtering, online content labeling, and offline labeling for Trust and Safety work. They are designed to interpret a user's written policy, which the announcement says allows generalization across products and use cases with minimal engineering. The models provide access to their reasoning process, which the announcement states facilitates debugging and trust in policy decisions, and users can configure reasoning effort at low, medium, or high levels based on use case and latency needs. The raw chain of thought is described as intended for developers and safety practitioners and not for exposure to general users or non-safety contexts.
OpenAI evaluated the models on internal and external evaluation sets. In the internal evaluation, multiple policies were provided simultaneously at inference time, and a test input was counted as accurate only if the model exactly matched golden set labels for all included policies. OpenAI also evaluated the models on the moderation dataset released with its 2022 research paper and on ToxicChat, a public benchmark based on user queries to an open-source chatbot.
Vinay Rao, CTO of ROOST, said the model is the first open source reasoning model with a "bring your own policies and definitions of harm" design, adding that in ROOST's testing it was skillful at understanding different policies, explaining its reasoning, and showing nuance in applying policies. ROOST is described as a non-profit established in 2025 that provides open source safety tools and technical support.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Ollama is partnering with OpenAI and ROOST (Robust Open Online Safety Tools) to bring the latest gpt-oss-safeguard reasoning models to users for safety classification tasks. gpt-oss-safeguard models are available in two sizes: 20B and 120B, and are permissively licensed under the Apache 2.0 license.