Alibaba just rolled out Qwen3.8-Max, and the company isn’t being shy about it. They’re calling it the largest and most capable AI model they’ve ever built. The timing isn’t random either—this launch puts Alibaba on a direct collision course with Moonshot AI, whose Kimi K3 model came out just last month, packing an even bigger parameter count.
The Numbers Game
Qwen3.8-Max runs on 2.4 trillion parameters. For context, parameters are basically the numerical settings a model tunes as it learns from data, helping it recognize patterns and actually perform tasks. Kimi K3, by comparison, holds 2.8 trillion parameters, so Moonshot still edges out Alibaba on raw scale.
That said, more parameters don’t automatically mean a better model. It’s more of a rough signal — an indicator of how much computing power and data went into building the system in the first place.
Chinese tech companies tend to be pretty open about these numbers, mostly because it helps attract developers. Most of their models are open-weight, meaning anyone can download and tweak them. That’s a different approach from what OpenAI, Anthropic, and Google do. Those companies keep their parameter counts under wraps for their closed-source models, for reasons they haven’t fully explained.
Read More: Alibaba Launches Qwen3: A New Line of Hybrid AI Reasoning Models
How It’s Performing So Far
Qwen3.8-Max made its debut on Arena.AI, a crowdsourced platform where people compare different AI models head-to-head. It landed at the top of the rankings among Chinese text-based models, which is a solid start.
It’s not beating everyone, though. The model still trails behind several Anthropic systems, including Claude Fable 5 and three separate Opus variants. There’s one bright spot worth mentioning: in the category that evaluates models handling images and other visual content, Qwen3.8-Max actually took second place globally. Only a Claude Fable 5 variant ranked higher.
What It Can Actually Do
Both Qwen3.8-Max and Kimi K3 can work across text, images, and video. Each one can also handle up to one million tokens in a single session, which is a pretty massive capacity.
Tokens, if you’re not familiar, are basically fragments of data — usually small pieces of words. The more tokens a model can process at once, the more it can handle in a single go. Think lengthy legal contracts, huge codebases, or documents running hundreds of pages, all processed without breaking them into chunks.
Under the Hood
Alibaba says the model runs on a mixture-of-experts setup. Basically, it doesn’t switch on the whole model for every request. Instead, it splits tasks between specialized components and only turns on whatever piece the job actually calls for.
That’s why just 95 billion parameters end up active at any one time. The full model’s way bigger than that. But this trimmed-down approach means lower running costs and faster replies, which starts to matter once you’re operating at real scale.
There’s also this—Alibaba pointed out that Qwen3.8-Max finished an entire software-engineering project in 16 days. They seem pretty proud of that one, holding it up as proof the model can actually deliver in the real world, not just on benchmarks. As for availability, developers won’t have to wait long. Qwen3.8-Max is expected to launch next week on Alibaba Cloud’s Model Studio platform.





