Mistral’s Trillion-Parameter Large 4 Tops Europe, But Chinese Rivals Still Win on Price

Sanket Chaukiyal

October 9, 2026

TL;DR

  • Mistral launched Large 4, a 1.05 trillion parameter mixture-of-experts model (49 to 52 billion active), in Research Public Preview on its API.
  • Independent evaluator Artificial Analysis scored it 38 on the Intelligence Index, the highest mark of any model built outside the US and China.
  • Chinese open-weight rivals Xiaomi MiMo-V2.6-Pro (46) and Z.ai GLM-5.3 (45) both score higher on the same index.
  • Running Large 4 costs over four times more per task than similar-intelligence Chinese models like GLM-5.3-Flash and DeepSeek V4.1 Flash; open weights are due at the end of October 2026.

Europe’s New Flagship Lands in Preview

Mistral AI just put its biggest model yet into the wild, sort of. Large 4 showed up this week in Research Public Preview on the company’s API, and it’s not a small step up from whatever came before. We’re talking 1.05 trillion total parameters, with only 49 billion to 52 billion of them active at any given moment thanks to a mixture-of-experts design.

That gap between total and active parameters matters. It means Mistral built something enormous without forcing every query to drag the entire model along for the ride. The company trained Large 4 on Nvidia’s Grace Blackwell GPUs, and the result handles both text and images across a 512k token context window, enough to swallow a novel and still have room for commentary.

Independent evaluator Artificial Analysis ran its own tests rather than taking Mistral’s word for it, and the numbers landed Large 4 at 38 on the firm’s Intelligence Index. That score makes it, in the firm’s own words, the most intelligent model from outside the US and China. On the security side, Large 4 pulled a 50% score on the Artificial Analysis Cyber Index, including an 82% result on the CyberGym-E2E-AA benchmark, a decent showing for a model still labeled a preview.

Pricing sits at $1.36 per million input tokens and $4.18 per million output tokens through the standard API tier. The weights themselves, the part that lets anyone actually download and run this thing, aren’t out yet. Mistral says they’re coming at the end of October 2026.

The Price Tag Problem

Here’s where the celebration gets complicated. Mistral built something genuinely impressive, the best model to come out of a country that isn’t the US or China, full stop. But “best outside two superpowers” is a category that exists mostly because those two superpowers have crowded everyone else out of the room.

Look at the actual leaderboard and Large 4’s 38 doesn’t even crack the top tier. Xiaomi’s MiMo-V2.6-Pro scores 46. Z.ai’s GLM-5.3 scores 45. Both are open-weight, both are Chinese, and both are already sitting above Mistral’s new flagship on the exact same Intelligence Index. Large 4 isn’t catching the frontier. It’s winning a regional contest while the actual race happens somewhere else.

And then there’s cost, which is where the story turns from awkward to genuinely rough. Running a task through Large 4 at standard pricing costs about $1.13. GLM-5.3-Flash does similar-intelligence work for $0.25. DeepSeek V4.1 Flash does it for $0.27. That’s not a rounding error. That’s Mistral charging more than four times as much for performance that Chinese labs are giving away at flash-tier prices.

Think of it like a boutique watchmaker proudly announcing the best mechanical watch made in Europe, only for the review to note that a factory down the road makes something just as accurate for a quarter of the price, in plastic, by the truckload. A boutique problem, basically. The craftsmanship isn’t fake. But craftsmanship doesn’t pay the bills when someone else is shipping the same accuracy at scale and at a discount.

I’ve watched enough model launches now to know that “best in region” headlines age badly the moment weights actually ship and developers start running real workloads against real budgets. If Large 4’s open-weight release in late October can’t close that pricing gap, the headline risks becoming a footnote within a quarter.

Is 38 on an intelligence index actually meaningful to anyone outside a handful of AI labs benchmarking against each other? Mostly, yes, in the sense that it tells enterprise buyers roughly where a model sits relative to competitors they’re already evaluating. But the gap between “highest score outside two countries” and “actually the model you’d choose” is where Mistral’s real test begins.

Mistral’s Long Game

Mistral AI has spent years building a reputation as the open-weight alternative when the conversation gets dominated by closed American labs and increasingly aggressive Chinese releases. The French startup made its name releasing models with published weights, letting researchers and companies actually inspect and run them rather than trust a black box behind an API.

That approach has earned Mistral a loyal following in Europe, where data sovereignty and model transparency carry real regulatory weight, not just marketing value. Large 4 fits that pattern, at least once the weights actually land. Until then it’s a preview, accessible through the API but not yet something a company can self-host or audit line by line.

The trillion-parameter scale itself is notable mostly because it shows Mistral can train at the size the biggest labs operate at, using the same class of Nvidia hardware everyone else is fighting over. That’s not a small technical achievement for a company that’s smaller than OpenAI, Google, or the major Chinese labs by almost any measure: headcount, funding, or compute budget.

But scale alone hasn’t been the deciding factor in this race for a while now. Chinese labs have been shipping competitive, cheaper open-weight models at a pace that’s made the “who trained the biggest thing” question almost beside the point. I’d want to see real deployment numbers before calling this a win for European AI, and I don’t think a benchmark score is the place to find them. Does training at trillion-parameter scale even matter anymore if a rival matches your score for a quarter of the cost?

What Comes After the Preview Label

The biggest open question is whether Large 4’s open weights, due at the end of October 2026, ship with the same pricing structure or whether Mistral adjusts once developers can actually self-host and compare real running costs against GLM and DeepSeek alternatives. Watch whether Mistral introduces a cheaper variant of its own, because that’s clearly the gap competitors are exploiting.

Also worth tracking: how Artificial Analysis and other independent evaluators rank Large 4 once the open weights are out and third parties can run their own benchmarks rather than relying on preview API numbers. Preview scores and post-release scores don’t always match.

Finally, keep an eye on whether European enterprise and government customers, the audience most likely to prioritize data sovereignty over raw price, actually adopt Large 4 at scale despite the cost gap. That adoption, or the lack of it, will say more about Mistral’s real position than any single benchmark score.

Editor's Note

What gets me about this one isn't the trillion parameters, it's the price tag. Mistral proved Europe can train at frontier scale, and I respect that. But claiming the best model outside the US and China feels thinner once you see Xiaomi and Z.ai beating it on the same index for a quarter of the cost. I'm watching the late October weights release closely. If Mistral can't match that pricing once developers self-host, this launch ages fast. Impressive engineering, shaky economics.

– Sanket Chaukiyal, founder, SmartChunks

FAQ

What is Mistral Large 4?

It's Mistral AI's new mixture-of-experts model with 1.05 trillion total parameters and 49 to 52 billion active parameters, trained on Nvidia Grace Blackwell GPUs and currently available in Research Public Preview on Mistral's API.

How does Large 4 compare to Chinese open-weight models?

Large 4 scored 38 on the Artificial Analysis Intelligence Index, the highest for any model outside the US and China, but Xiaomi's MiMo-V2.6-Pro (46) and Z.ai's GLM-5.3 (45) both score higher overall.

Why is Large 4 so much more expensive to run?

At standard pricing of $1.36 and $4.18 per million input and output tokens, a typical task costs around $1.13, more than four times what similar-intelligence models like GLM-5.3-Flash ($0.25) and DeepSeek V4.1 Flash ($0.27) charge.

When will Large 4's open weights be available?

Mistral has said the open weights are planned for release at the end of October 2026, though the model remains in preview on the API until then.


Source: Tom's Hardware

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → • LinkedIn