Mistral OCR 4.1 aims to make scanned documents easier for AI to read

Mistral has introduced Mistral OCR 4.1, a public preview service designed to turn documents into machine-readable content while retaining information about their layout. The company announces in Mistral Docs that the model can identify paragraphs, assign labels to structural blocks and return confidence scores for each block.

The model, called mistral-ocr-4-1, is part of Mistral’s Document AI offering. It is intended for workflows that need more than plain extracted text. A document may contain headings, paragraphs, tables or other sections. Knowing where these elements appear on a page can help software organise content for review, search, archiving or later processing by AI systems.

Layout data alongside extracted text

OCR, short for optical character recognition, converts text in scans, photographs and PDFs into digital text. Mistral OCR 4.1 adds paragraph-level bounding boxes. These are coordinates that mark the position and size of a paragraph on the original page.

According to the documentation, the service also provides structural block labels and block-level confidence scores. Labels describe what type of content the system has found. Confidence scores indicate how certain the model is about a result. That can help teams identify passages that may require human verification, especially in large document collections.

Mistral makes the service available through its /v1/ocr endpoint for OCR, bounding box extraction and structured annotations. It also supports batch processing through /v1/batch, allowing organisations to submit larger groups of documents.

The listed price is €3.50 per 1,000 pages. Mistral lists annotated pages at €4.38 per 1,000 pages. The company describes the release as Premier v4.1 and a public preview.

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×