Qwen 3.8 Max: Everything You Need to Know About Alibaba’s Latest AI Model
By Veer Solanki · · 1489 words
Topics: AI, 2.4 Trillion Parameters, AI APIs, AI Benchmarking, AI Coding, AI Ecosystem, AI Hardware, AI industry
Monday morning saw another release from Alibaba. This one is even bigger than the previous one – 2.4 trillion parameters. For output, $2 per million input tokens, $6 compared to $0.25 for cached reads. The model it replaces, Qwen 3.7 Max, cost $2.50 and $7.50, respectively.
It gets bigger, better and cheaper every couple of weeks or so – no objections there.
What stands out is the second sentence. Weights are coming – next week, according to the announcement. Weights for the flagship, not the distilled or community versions. At the same time, there’s a 27B sibling that will also be open.
For three generations, Alibaba has had a simple policy: the open versions (Max and above) have weights, the ones below do not, and the flagship was always behind an API key. The open line had a peak at 3-235B, but that was never the best model Alibaba had. And now, this neatly organised pricing seems to have been dismantled – and few are likely to cry over the loss of the $2500 monthly subscription, I imagine.
The pitch is not a chatbot
Looking at the announcement, a surprising discovery was made – the actual use cases take far more space than the general capabilities. There is the usual “unattended coding”, but it appears to have been supercharged – “coding a self-evolving harness from an empty folder across ten-plus days of unattended operation.” The other claims include 500 turns of chip design optimisation, e-commerce simulations over 365 days, and an autonomous research loop. That one self-explains but is notable for spending “one hundred and twenty-five hours on reconstructing the pipeline of a paper and innovating a data selection method that excels the cited work by 2.71 points.”
The benchmarks follow the same pattern. Not one of them is about question-answering. All are about training loops or pipelines, with the sole focus on “retention of the intended operations in the ninety-ninth hour.
It is a model intended for long-running trainers, and the marketing knows it. Even the multimodal aspect is about training – vision is downplayed as a part of the “loop,” focused on planning and self-correction rather than content creation.
The numbers, and who ran them
The claims are the vendor claims, of course. There are plenty of impressive-sounding statements, such as the design optimisation loop cutting the number of gates from 8,298 to 678 while dropping the die size by 81% without dropping the clock speed below 500 MHz, or the 4.16x headroom achieved in a 365-day e-commerce simulation. The runs at the WWW2025 Multimodal Dialogue Intent Recognition Challenge, with a 13-percentile finish among 526 human competitors within twenty-four hours of launch, are remarkable as well.

All are from Alibaba, using their training loops, and the usual caveats apply – no reason to doubt their word, no reason to believe their numbers. But the third-party results were out hours later, and that is the point.
Vals AI put the Qwen 3.8 Max at 66.1 on the Vals Index, second in the open category and tenth overall among the 43 models on their testing list. That is an exact tie with Claude Opus 4.7’s score on the same index but at 2.3x lower price (this one at $2.68 against $6.17). On the SWE-bench, it is at 87.3%, beating GPT-5.5’s 82.6% and GLM-5.2’s 83.3%, but below Claude Opus 4.8’s 89.2%. On Terminal-Bench 2.1, it achieved a 67.4, better than the previous Max’s 61.0.
But vals added an important footnote to the last result, which is worth considering. The Alibaba results vs. the Terminal-Bench 2.1 timeout rules have modified timeouts. Vals left theirs as stated in the benchmark, meaning that the same Qwen 3.8 Max would have scored 66.4 on the same clock as all the other models tested.
Arena has a few relevant results as well – the Qwen 3.8 Max hit the Frontend Code Arena as number four with a 1,668 Elo rating, just behind Claude Opus 5’s 1,705 and Kimi K3’s 1,676, and tied with Opus 5 at max effort with a 1,669 score. It took second in the Vision Arena with a 1,305, thirteen points below Claude Fable 5.
(Vals are always fun to look at but not worth overthinking – taking the 3.8 Max’s 66.1 on the Vals Index, compare it to the 57.5 of the previous Max of the same generation (3.7) – that’s an increase of 8.6 over ten weeks, along with a price cut.) This is the trajectory, not the absolute value. The one worth noting is the one everyone will use as a point of reference.)
What “open” is doing in that sentence
Even before the launch, but a few hours into it, some people started looking at the licence.
Ostris noted that certain jurisdictions, including the USA, the EU, the UK and South Korea, are explicitly excluded from use, suggesting that downloading the weights from these locations may be illegal. Alibaba has yet to comment officially, and as of this writing, the licence remains as is.
The same arguments were made about MiniMax H3, which initially appeared to have a similar regional restriction. In that case, it was later clarified that these jurisdictions simply require an explicit authorisation for the use of the weights, which is a genuine difference. Which, in turn, is why the ambiguity is worth consideration.
Open weights can be the weights you download, and they can be something else entirely. They can be a statement without OSI rights, use-case restrictions, jurisdictional limitations or even a lack of commercial permissions in the place you operate from. The phrase has stretched considerably over the years, and no one in a launch post is particularly inclined to parse the fine print.
The other thing you cannot do with it
Let’s say the licence is fine.
You still cannot run it.

2.4T mix-of-experts only activates a fraction of itself per request, presumably around 95 billion parameters or 4%, and that is where the price decrease comes from – but not the memory requirements. Jamin Ball estimates that this class of model requires over a terabyte of memory just to load the weights, with no less than eight H100S or B200S required to serve it. And that is before getting to the supernode recommendations of the highest-end cards.)
In short, the flagship model is available for open download but not personal use, for almost everyone, and certainly for almost all practical purposes.
Which is why Qwen3.8-27B may well end up being the more significant release. The 27B variant is the one that will run on people’s home computers, and if even a fraction of the flagship’s capabilities make it into that version, that is where people will start adopting it. The 2.4T model is a legitimising force – an anchor for the 27B’s claims and a signal to the market that Alibaba has something worth serious consideration -, but it will not install on anyone’s PC.
What this actually means
I wrote about Kimi K3 last week, and this class of release is similar to that one. In fact, it exemplifies it – Kimi K3 is a 2.8T model, and there is also GLM-5.2, the DeepSeek V4 Flash, the MiniMax H3 – and now the Qwen.
Artificial Analysis has generally scored the Chinese frontier models as being between three and nine months behind the best American closed weights. I think that is right, but the open frontier has also been Chinese for about two years, and this is not up for debate anymore. Artificial Analysis’s front-end design rankings have two Chinese models in the top three, and the vision capability gap has been narrowed to thirteen Elo points over Fable 5.
The strategic reason for Alibaba’s choice to keep Max closed was obvious – a simultaneous release would make their weights no longer exclusive, and the price difference between their offering and the open alternatives was simply too small. The value of the API key was eroding – and it was no longer worth it, particularly with several similar weights becoming available. And now, for the first time, they have chosen to move towards openness and ecosystem building.
I think this is generally a good sign. Lower prices, more weights, and a serious reversal of Alibaba’s closed stance from last year are all positive developments.
But there is another reading of this event, and it is not as cheerful. The open frontier model is one you are not licensed to download, that you cannot run even if you do, promoted by a company that benefits immensely from you using their API in either case. And this is the moment when openness became a distribution channel for influence over the entire stack. One where the choice between using an open model and an API is no longer a technical or economic consideration but a geopolitical one at the system level.
Both readings are correct. Only the first one is not the headline.