Mistral unveils Large 4, a trillion-parameter open-weight AI model

Mistral has opened a public preview of Mistral Large 4, its largest model so far and a major addition to the market for open-weight AI systems. The company says it will release the model weights later this month, allowing organizations to run and adapt the system on their own infrastructure.

The French AI company describes Mistral Large 4, also called ML4 or “Le Chonk,” as a natively multimodal model. It can process text and images and is designed for coding, AI agents, business workflows, document analysis, and scientific work. Mistral says the model has 1.05 trillion total parameters, although its Mixture-of-Experts design activates 49 billion parameters for a given task. Its documentation lists a context window of up to one million tokens.

The release matters because the largest high-performing models are usually controlled by their providers. Open-weight models give customers the option to inspect, customize, and deploy a model in a private cloud or their own data center. This can be important for organizations with strict data, compliance, or operational requirements.

Focus on enterprise use and cybersecurity

Mistral says it trained ML4 from scratch in its European data centers using 3,800 Nvidia Grace Blackwell GPUs. The company positions the model as a sovereign option for customers that want to retain control over how and where an AI system runs. It says a European deployment will operate under European law and independently from other digital service providers.

Cybersecurity is central to that strategy. Mistral argues that provider-level safety restrictions in closed systems can interfere with legitimate security research, incident response, and vulnerability testing. The company says ML4 can be deployed under an organization’s own policies, while still refusing malicious cyber requests at a higher average rate than the open models it compared.

According to Mistral, ML4 ranks among the five strongest models worldwide on the Artificial Analysis Cyber Index. It reports an 82 percent score on one vulnerability reproduction and repair test, plus a 93 percent result on Cybench, a set of cybersecurity exercises. These are company-reported benchmark results and should be interpreted alongside independent testing once the weights become available.

The model’s open-weight status does not mean it will be easy to operate locally. As Sabrina Ortiz writes for The Deep View, a model of this size will be beyond the reach of most desktop users and may also be difficult for many universities to host. Mistral is therefore offering the preview through its API while targeting enterprise deployments in private cloud and on-premises environments.

Claims across coding, agents, and vision

Mistral reports strong results in agentic coding, a category where models use tools such as terminals and code repositories to complete multi-step software tasks. It says ML4 scores 61.7 percent on DeepSWE v1.1 and 28.3 percent on Terminal-Bench 4. The company also reports that professional reviewers in a blind evaluation ranked its preview model second among five systems, behind Claude Opus 5.

For general office work, Mistral highlights AutomationBench, which tests workflows across software such as Gmail, Google Sheets, Slack, and Salesforce. ML4 scored 59.9 percent, according to the company. It also claims improvements in producing spreadsheets, presentations, PDFs, and legal or financial analyses.

Vision is another focus. Mistral says ML4 can inspect technical drawings, retrieve evidence from PDFs, and analyze large satellite images. It reports that the model exceeded GPT-6 Astra on Dense 200, a visual grounding benchmark, by one percentage point.

The API preview supports structured outputs, function calling, document question answering, batching, agents, and built-in tools. Mistral lists pricing of $0.68 per million input tokens and $2.09 per million output tokens, with lower rates for cached input.

The company says its reinforcement learning work is continuing and that the preview model may improve before the planned weight release. For customers, the key question will be whether independent evaluations confirm the stated results and whether ML4’s scale can be managed at a cost that makes its autonomy attractive.

Sources

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×