Alibaba Open-Sources Qwen3.8-2.4T-A95B as First Qwen-Max Flagship

Sanket Chaukiyal

August 16, 2026

TL;DR

  • Alibaba’s Qwen team has released full open weights for Qwen3.8-2.4T-A95B, calling it the first Qwen-Max-class flagship model ever fully open-sourced.
  • It’s a Mixture-of-Experts model with 2.4 trillion total parameters but only about 95 billion active per pass, spread across 512 experts per layer (10 routed, 1 shared).
  • Native context runs 262K tokens, extensible to 1 million, and the weights total 4.9TB in BF16 format under a custom Qwen3.8-Max license.
  • Self-reported scores hit 86.6 on Terminal-Bench 2.1, 93.0 on PaperBench, 92.6 on GPQA Diamond, and 67.7 on SWE-bench Pro.

What Alibaba Actually Shipped

Qwen-Max has always been the closed, top-shelf option in Alibaba‘s model lineup, the one you paid per token to access through an API and never actually held. That changed this week. The Qwen team published full weights for Qwen3.8-2.4T-A95B on Hugging Face and ModelScope, and the accompanying note doesn’t hedge: this is, in their words, “the first time a Qwen-Max-class (flagship) model has been fully open-sourced.”

The architecture is the interesting part. It’s a fine-grained MoE design with 2.4 trillion total parameters, but a forward pass only activates roughly 95 billion of them, routed through 512 experts per layer, ten selected plus one shared expert that’s always on. Layer that on top of a hybrid Gated-DeltaNet and Gated-Attention setup, and you get a model built to scale without every parameter firing at once.

Context handling is where it gets genuinely useful for real workloads. Native context sits at 262K tokens, extendable to a full million, with text-only generation and reasoning mode forced on by default. And the weights themselves aren’t small: 4.9TB in BF16, published under a custom license rather than a standard permissive one like Apache 2.0.

Our Take

Downloading this thing isn’t like grabbing an app off a store shelf. It’s closer to Alibaba backing a moving truck up to your driveway, unloading 4.9 terabytes of finished machinery onto your lawn, and handing you the keys before driving off. You own it now. Good luck finding a garage big enough to park it in.

That’s the tension nobody’s talking about enough. Open weights sound like a democratizing move, and in one sense they are: teams that once had no choice but to pay per-token API fees for flagship-tier reasoning now have a downloadable alternative. But running a 2.4-trillion-parameter model, even a sparse one, requires GPU memory most individual developers and plenty of small companies simply don’t have sitting around. Open doesn’t automatically mean accessible.

I’ve watched enough open-weight drops promise to “close the gap” between proprietary and open models that I’ve grown a little numb to the phrase. This one comes with actual numbers attached, though, and the fact pack’s own framing puts it plainly: this closes the open-versus-closed gap at the very top of the capability curve. Whether that holds up once independent labs run their own evaluations is a separate question from whether Alibaba believes it.

Worth asking: does releasing a flagship model as open weights actually threaten the closed API business model, or does it just create a two-tier market where only well-resourced teams can afford to self-host the good stuff? Right now it looks like the latter. Not a small shift, though. A real one.

From Paid API to Free Download

Qwen-Max previously existed only as a commercial API product, the kind of model Alibaba sold access to rather than gave away. Qwen3.8-2.4T-A95B changes that calculus for one specific model tier, not the entire lineup, but it’s the first time the top tier has crossed over into open weights at all.

That matters for the economics of self-hosting and fine-tuning specifically. Teams that were locked into metered API pricing because no open alternative matched flagship capability now have a legitimate option to evaluate, assuming they can secure the infrastructure. Fine-tuning a model this size isn’t a weekend project on a laptop, but for organizations already running serious GPU clusters, it removes a real constraint: dependency on a single vendor’s API terms and pricing.

The benchmark numbers, self-reported as they are, land in territory that used to be exclusive to closed frontier models. Terminal-Bench 2.1 at 86.6 and PaperBench at 93.0 aren’t token gestures. GPQA Diamond at 92.6 and SWE-bench Pro at 67.7 suggest a model built for genuinely hard reasoning and coding tasks, not just chat. None of that has been independently verified yet, and that caveat matters more than usual given how large a claim “first fully open Qwen-Max flagship” actually is.

Signals Worth Tracking

The next few weeks will tell you more than Alibaba’s own announcement can. Watch for independent benchmark replications from labs and research groups outside Alibaba’s own testing pipeline, since every number in this release so far is self-reported. Watch, too, for how fast the open-source community produces quantized or distilled versions that fit on hardware smaller than a data center, because that’s what will actually determine whether this model reaches beyond a handful of well-funded labs. And keep an eye on whether competing labs respond by open-sourcing their own top-tier models rather than keeping them gated behind paid APIs, since that’s the real test of whether this release shifts the broader market rather than just adding one more entry to it.

Editor's Note

What gets me here isn't the parameter count, it's the license fine print. Alibaba is calling this fully open, but a 4.9TB model under a custom license isn't the same thing as Apache 2.0, and I think people are glossing over that distinction. I want to see who actually runs this beyond a handful of labs with spare GPU clusters. If independent benchmarks match Alibaba's self-reported numbers, that's the real story here, not the release headline itself.

— Sanket Chaukiyal, founder, SmartChunks

FAQ

What is Qwen3.8-2.4T-A95B?

It's a Mixture-of-Experts language model from Alibaba's Qwen team, with 2.4 trillion total parameters and roughly 95 billion active per forward pass. Alibaba describes it as the first Qwen-Max-class flagship model to be released as full open weights, rather than kept behind a paid API.

How big are the actual model files?

The published weights total 4.9TB in BF16 format, hosted on Hugging Face and ModelScope. That size reflects the 2.4 trillion parameter count and means running or downloading the model requires substantial storage and GPU memory, not something most individual developers can casually spin up.

What license covers the weights?

The model is published under a custom Qwen3.8-Max license rather than a standard permissive open-source license like Apache 2.0 or MIT. Anyone planning to deploy or fine-tune it commercially should read that license directly rather than assume it behaves like a typical open-source release.

Are the benchmark scores independently verified?

No. The scores, including 86.6 on Terminal-Bench 2.1, 93.0 on PaperBench, 92.6 on GPQA Diamond, and 67.7 on SWE-bench Pro, come from Alibaba's own testing. Independent replication from outside labs hasn't been reported yet, so those numbers should be treated as self-reported until third parties confirm them.


Source: Hugging Face

Sanket Chaukiyal — Editor at Smart Chunks

Sanket Chaukiyal

Technology editor • 12+ years in editorial

Sanket is the founder and editor of Smart Chunks. He spent over six years at Autocar India (Haymarket SAC Publishing) as Sub Editor and Senior Copy Editor, and later served as Account Director (Content) at Rite Knowledge Labs. He holds a Master's in Media and Communication from the Symbiosis Institute of Media and Communication.

All articles → LinkedIn