AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Introducing Mistral Small 4

Collected Oct 1, 2026

Mistral AI announced Mistral Small 4, the next major release in its Mistral Small family. According to the company, it is the first Mistral model to combine the capabilities of Magistral for reasoning, Pixtral for multimodal input, and Devstral for agentic coding into one model. It is released under the Apache 2.0 license.

The model is a hybrid designed for general chat, coding, agentic tasks, and complex reasoning, and accepts both text and image inputs. Mistral AI said it is joining the NVIDIA Nemotron Coalition as a founding member.

Architecturally, Mistral Small 4 uses a Mixture of Experts design with 128 experts and 4 active per token. It has 119B total parameters, with 6B active parameters per token and 8B including embedding and output layers. It supports a 256k context window. A reasoning_effort parameter allows toggling between fast responses and deeper reasoning; Mistral said reasoning_effort="none" is equivalent to the chat style of Mistral Small 3.2, while reasoning_effort="high" has verbosity equivalent to previous Magistral models.

Mistral reported a 40% reduction in end-to-end completion time in a latency-optimized setup and 3x more requests per second in a throughput-optimized setup compared to Mistral Small 3. Minimum infrastructure is listed as 4x NVIDIA HGX H100, 2x NVIDIA HGX H200, or 1x NVIDIA DGX B200, with recommended setups of 4x H100, 4x H200, or 2x DGX B200.

The model is available through the Mistral API and AI Studio, and on Hugging Face. Mistral said it is available on vLLM, llama.cpp, SGLang, and Transformers, and is available day-0 as an NVIDIA NIM. Pricing listed is $0.15 per million input tokens and $0.6 per million output tokens.

Read at Mistral AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt