Mistral AI Unveils Trillion-Parameter Model as European Challenger Takes Aim at AI Leaders

Mistral AI Unveils Trillion-Parameter Model as European Challenger Takes Aim at AI Leaders

Image to help understand the article

Mistral Large 4 Debuts in Preview

French artificial intelligence company Mistral AI unveiled Mistral Large 4 on Oct. 6, 2026, presenting the trillion-parameter system as a new contender in coding, workplace automation, cybersecurity and visual reasoning.

Nicknamed “le Chonk,” the model uses a mixture-of-experts architecture, activating 49 billion of its 1 trillion parameters for a given task. It supports more than 160 languages and can process multimodal inputs while producing text-only output.

The model, identified as mistral-large-4-0, is initially available through a preview API in Mistral Studio. Pricing is $1.36 per million input tokens and $4.18 per million output tokens.

Model Weights to Follow Safety Testing

Mistral did not release the model’s weights at launch. The company said it plans to make them available Oct. 27 after roughly three weeks of testing with trusted partners and governments. Until then, users can access the model only through public-preview and guardrail-protected endpoints.

The company said the staged release is intended to support defensive uses, particularly in cybersecurity, while limiting the model’s potential use in malicious attacks. The weights will be distributed under a custom Mistral license.

Mistral Large 4 was trained for about two months using 4,000 Nvidia Grace Blackwell graphics processing units. Mistral said reinforcement learning for the preview remains underway and that the model has not shown signs of reaching its performance ceiling.

Strong Company Results, but Mixed Outside Testing

Mistral reported a 61.7% score on DeepSWE v1.1 and said the model outperformed DeepSeek V4 Pro 0813. It also posted 28.3% on Terminal Bench 4.0, 59.4% on SWE-Atlas-QnA and a combined Coding Agent Index score of 49.8%, according to the company.

Results were less favorable in a blind human coding evaluation by Surge AI. Mistral Large 4 received 3.74 out of 5, compared with 4.22 for Anthropic’s Claude Opus 5. Another DeepSWE tally placed Mistral at 62%, ahead of several Chinese-developed models but behind GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5, which each scored about 74% under their best publicly reported configurations.

Mistral also reported a 59.9% score on AutomationBench and an Elo rating of 1,393 on AA-Briefcase, which measures knowledge-work performance. Its results included 15% on Harvey Legal Agent and 67% on Finch Finance, tying DeepSeek V4 Pro 0813 on the latter test.

Cybersecurity and Visual Reasoning

The company said Mistral Large 4 ranked among the world’s top five models on the AA Cyber Index and scored 93% on Cybench, surpassing major open-weight competitors. It reported an attack-resistance score of 93.3% on the B3 Security Benchmark and 1.691 out of 2 on the KORA Benchmark’s responsible-response evaluation.

In visual localization tests, Mistral reported a 42% score on Dense 200, narrowly above GPT-6 Astra’s 41%, and 73% on DIOR-RSVG. The company also described the model as a leading open-weight performer on SciCode-Verified.

CEO Arthur Mensch said the model was ahead of Chinese models in some areas, including cybersecurity, but did not specify which models or comparison methods he meant.

European Model Still Trails Top U.S. Rivals Overall

Independent evaluations presented a more restrained picture. Vals AI gave Mistral Large 4 a Vals Index score of 48.05%, ranking it 32nd among 44 models. Its strongest placement there was sixth among 75 models on Harvey Legal Agent, while its rankings were lower on finance, tax, business and coding evaluations.

Artificial Analysis gave the model an Intelligence Index score of 38.4, placing it eighth among open-weight models. That was above DeepSeek V4 Pro and South Korea’s Motif 3 but below several Chinese-developed systems. It also remained well behind leading closed models from U.S. companies, including Claude Opus 5.5 at 57.6 and GPT-6 Astra at 52.7, as well as Gemini 4 Argon at 52.6.

Artificial Analysis estimated Mistral Large 4’s cost per task at $1.13, or 1 euro. Mistral science Vice President Pierre Stock said the model was trained with substantially less computing power than closed competitors and roughly two to three times less than Chinese rivals, though some benchmark results and exact competitor scores have not yet been released.

The launch follows Mistral’s September fundraising round, when it raised 3 billion euros at a valuation of 21 billion euros.

Source: Original Korean article - Trendy News Korea

Comments