AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Introducing Mistral OCR 3

Collected Oct 5, 2026

Mistral AI has introduced Mistral OCR 3, a document extraction model that the company says delivers a 74% overall win rate over Mistral OCR 2 on forms, scanned documents, complex tables and handwriting. Mistral described the model as state-of-the-art accuracy, outperforming both enterprise document processing solutions and AI-native OCR solutions.

The model extracts text and embedded images from a wide range of documents and supports markdown output with HTML-based table reconstruction, including colspan and rowspan tags to preserve layout. Mistral said it is a much smaller model than most competitive solutions and is priced at $2 per 1,000 pages, with a 50% Batch-API discount reducing the cost to $1 per 1,000 pages. Annotations are listed at $3 per 1,000 pages.

Mistral OCR 3 is available through the API under the model name mistral-ocr-2512 and through Document AI, a UI in Mistral AI Studio that parses PDFs and images into clean text or structured JSON. Mistral said the model now powers Document AI Playground in Mistral AI Studio. It is fully backward compatible with Mistral OCR 2, and a self-hosting option is offered for organizations with stringent data privacy requirements.

Mistral said it introduced more challenging internal benchmarks based on real business use-case examples from customers, comparing model outputs to ground truth using a fuzzy-match metric for accuracy. The company said the model is a significant upgrade across all languages and document form factors compared with Mistral OCR 2.

Use cases listed by Mistral include extracting text and images into markdown for downstream agents and knowledge systems, automated parsing of forms, invoices and operational documents, end-to-end document understanding pipelines, and digitization of handwritten or historical documents. Mistral said early customers are using the model to process invoices into structured fields, digitize company archives, extract clean text from technical and scientific reports, and improve enterprise search.

"OCR remains foundational for enabling generative AI and agentic AI," said Tim Law, IDC Director of Research for AI and Automation. "Those organizations that can efficiently and cost-effectively extract text and embedded images with high fidelity will unlock value and will gain a competitive advantage from their data by providing richer context."

Read at Mistral AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt