Introducing Shieldstral.
Mistral AI announced Shieldstral, a 3B open-weights multimodal safety classifier released under the Apache 2.0 license. According to Mistral, Shieldstral matches or outperforms open guard models up to 7 times its size on text safety and sets a new state of the art on multimodal moderation. The company says the model runs on a single 16GB NVIDIA GPU.
Shieldstral frames content moderation as a binary question-answering task with three parts: an instruct section giving evaluation context, strictness, and optionally a definition of unsafe content; a query containing a single yes/no question; and a document holding the content to judge, which may be a prompt, a response, a prompt-response pair, or an image with optional text. At inference, Mistral says, the model reads only the yes and no logits and softmax-normalizes them into a continuous safety score.
Mistral says the approach unifies prompt classification, response moderation, refusal detection, and toxicity detection, and lets policies live entirely in the prompt so one checkpoint adapts to new policies at deployment time without retraining. The company says Shieldstral was trained on real and synthetic data with diverse label formats and taxonomies consolidated into one framework.
Mistral describes building the model by converting heterogeneous public safety datasets into a shared instruction-query-document format, calibrating strictness per source, constructing contrastive policy pairs generated by an LLM, supplementing scarce visual safety data with general-purpose image datasets filtered through a vision-language reranker, and merging complementary checkpoints via LoRA fine-tuning and SLERP. It says the model was built end to end on Forge, Mistral's platform for training, aligning, and evaluating custom models. Mistral states that all evaluation samples were held out from training.
Mistral also said it is an inaugural member of the Open Secure AI Alliance with NVIDIA and other organizations. It said it is continuing work on multilingual coverage, longer-document robustness, and broader multimodal safety. Weights are available for download; documentation is offered for text and image safety classification against a supplied policy.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.