French lab Mistral put Large 4 Preview on its API. A third-party benchmark firm then put a number on it—and framed that number as a geographic comeback story.
On 6 October 2026, Unite.AI reported Artificial Analysis’s analysis of Mistral Large 4 Preview: a score of 38 on Artificial Analysis Intelligence Index version 4.3.2. [1]
What the third-party score actually is
Artificial Analysis is a third-party benchmark firm. Its Intelligence Index v4.3.2 is described as a weighted average of ten evaluations across four categories: agents 30%, coding 20%, scientific reasoning 20%, and general 30%. The firm estimates a 95% confidence interval of less than ±1% and says the suite is primarily text-based and English-language. [1]
On that index, Unite.AI says Artificial Analysis scored Mistral Large 4 Preview at 38—comparable, in the firm’s framing, to GPT-6 Luna (max) at 38 and DeepSeek V4.1 Flash (max) at 39—and described it as the most intelligent model from outside the US and China. On Artificial Analysis’s model page, that score is reported as ranking the preview 64 of 225 models in its comparison class, above a class median of 26. [1]
Those rankings and peer comparisons are third-party benchmark results, not Mistral’s own marketing scores.
Other Artificial Analysis metrics
According to the same Unite.AI write-up of Artificial Analysis:
- Cyber Index: 50 (level with GLM-5.3-Flash; behind MiMo-V2.6-Pro at 56).
- CyberGym-E2E-AA: 82%.
- GDP.pdf (document reasoning): 19%—described as an 18-point gain versus Mistral Large 3, partly tied to an API change allowing 100 images per request (up from eight).
- Cost to run the Intelligence Index: about $1.13 per task at list prices ($1.36 / $4.18 per 1M input/output tokens), or about $0.57 with a reported 50% two-week launch discount.
- Output speed via Mistral’s API: 116.1 tokens per second; time to first token 1.46 seconds.
- Context window listed: 512k tokens; model described as 1 trillion parameters with 49 billion active. [1]
Artificial Analysis also characterises the preview as more than four times the cost per task of similar-intelligence open-weights models, and notes weights are not yet public, so the preview is currently classified proprietary. [1]
What Mistral claims (separately)
Company claims, as relayed by Unite.AI: public preview launched 6 October 2026 (nickname “Le Chonk”); natively multimodal; trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s European datacentres; training data spanning more than 160 languages including every official EU language; open weights planned by end of October 2026; reinforcement learning for the preview still in flight; and a claim that the model significantly outperforms any open-weight model developed in the US or Europe. [1]
Those are company assertions. They are not the same thing as the Artificial Analysis score of 38.
What we don’t know / What this does not prove
- An Intelligence Index score is one firm’s weighted suite—not proof of real-world usefulness, safety or multilingual quality.
- Peer model names and scores above come from Artificial Analysis via Unite.AI; this draft does not re-run the benchmarks.
- Open weights “by end of October” remain a company plan, not a completed release.
- Cost figures use list/discount token prices Artificial Analysis applied; billable spend for a given customer may differ.
The Bottom Line
Third-party firm Artificial Analysis scored Mistral Large 4 Preview at 38 on Intelligence Index v4.3.2, with separate cyber and document-reasoning figures and a relatively high measured cost per task. Treat Mistral’s training, language and open-weight promises as company claims until weights and independent checks catch up.
Sources
- Unite.AI (Jonas Reeve; AI-generated, editorially reviewed per Unite.AI disclosure), “Benchmark Firm Gives Mistral Large 4 Preview a 38 Intelligence Score,” 6 October 2026 — https://www.unite.ai/benchmark-firm-gives-mistral-large-4-preview-a-38-intelligence-score/

