The French AI company trained ML4 from scratch over roughly two months using 4,000 Nvidia Grace Blackwell GPUs in its European data centers. Mistral says the training covered more than 160 languages, including all official European Union languages.
The model is currently available in preview through Mistral's API. Mistral plans to release the model weights on Oct. 27 following about three weeks of testing involving developers, cybersecurity executives, and government authorities. The company also plans to continue reinforcement learning during the preview before releasing the final checkpoint.
“ML4 is at the frontier of open weight models,” Mistral cofounder and chief scientist Guillaume Lample said in materials shared with VentureBeat.
Mistral is positioning the model as an alternative for organizations that want more direct control over their AI infrastructure. Once the weights are available, customers will be able to customize and operate the model themselves, including on sovereign infrastructure. Mistral also says deployments can support zero-data-retention configurations.
The company is emphasizing specialized capabilities alongside general model performance. ML4 is being targeted at software engineering, cyberdefense, financial analysis, satellite and aerial imagery, technical drawings, chip design, and other technical applications.
“There are a lot of areas where the other labs will not focus that much,” Lample told WIRED. “There are so many domains in which you can improve models.”
Cybersecurity is one of the most prominent parts of Mistral's pitch. The company argues that organizations using closed AI services for defensive security work can be affected by provider restrictions on dual-use requests. With an open-weight model, businesses can instead configure the system for code scanning, defensive testing, and other security workloads without relying entirely on a third party's moderation policies.
That argument has taken on greater significance as access to powerful AI systems becomes more contested. The Trump administration imposed temporary restrictions in June on distributing models from OpenAI and Anthropic because of concerns about their possible use in advanced cyberattacks. The White House has also reportedly asked U.S. AI companies to withhold unreleased models from the UK's AI Safety Institute until they undergo U.S. review.
Mistral is using its open-weight approach to differentiate itself from those closed systems as well as from models developed in China. The company describes ML4 as the strongest open-weight model created outside China and says its performance is approaching some proprietary systems. Mistral also says the model was trained from scratch rather than relying on distillation from a larger competitor.
“Mistral is still in the race of getting the best model,” Lample told WIRED. “This is the main message.”
Open-weight models can also give businesses a different cost structure because organizations can run the models on their own compute instead of paying a provider for every interaction. Mistral generates revenue through usage-based access to its cloud services and by providing engineers who help customers adapt its models for particular applications.
The company has historically operated with fewer financial and computing resources than OpenAI and Anthropic and has trailed the larger U.S. labs in areas including model performance and release frequency. Mistral has recently expanded its resources, however. In September, the company raised $3.3 billion at a $24 billion valuation, while its earnings have reportedly increased twentyfold over roughly the past year.
Lample has also expanded Mistral's science team from three researchers to about 300, according to the company. He said Mistral expects ML4's capabilities to continue improving as reinforcement learning progresses and its training infrastructure grows.
The Le Chonk name also reflects an online joke that developed around expectations for a much larger Mistral model. A fictional system called “Le Chaton Fat” circulated on X and Reddit earlier this year with exaggerated specifications for a massive French AI model. Mistral executives joined in on the joke, and the company later adopted another oversized-cat reference as ML4's internal nickname.
Behind the name is Mistral's attempt to demonstrate that an independently deployable model can remain competitive with much larger proprietary systems. By combining a trillion-parameter architecture with downloadable weights and a focus on specialized technical workloads, the company is betting that control over where and how a model runs can become as important to some customers as raw model performance.
This analysis is based on reporting from Wired.
Image courtesy of Deeploy.
This article was generated with AI assistance and reviewed for accuracy and quality.