CodeBucks logo
WangDou

Alibaba Launches Qwen3.8-Max: 2.4 Trillion Parameters, Beating GPT-5.6 Sol and Fable 5 on Multiple Benchmarks

2026-08-05·WangDou AI Express·Alibaba / Qwen / LLM

Alibaba just slapped a scorecard on Silicon Valley's desk — Qwen3.8-Max outperforms OpenAI's and Anthropic's flagship models on multiple benchmarks.

Three Key Takeaways

2.4 trillion parameters, 95 billion active. Qwen3.8-Max is the largest model in Alibaba's Qwen family, packing 2.4 trillion total parameters in a mixture-of-experts architecture with roughly 95 billion active at any given time. This design balances inference cost against model capability — you do not need to light up every parameter to get frontier-level performance.

Leading scores on key benchmarks. IFBench (instruction following): 82.8, versus 72.7 for GPT-5.6 Sol and 63.5 for Fable 5. PaperBench (paper reproduction): 93.0, ahead of Sol's 90.5 and Fable 5's 88.8. HealthBench: 60.2 and PLawBench: 73.2, also in the lead. On Terminal Bench 2.1 it trails Sol slightly at 86.6 versus 88.8, but pulls ahead on AndroidBench (75.1 vs 74.0) and Alibaba's own QwenSWEBench (80.7 vs 73.5). All benchmark data is as of August 3; no independent third-party evaluations have been published yet.

A wave of Chinese model launches is intensifying the price war. Bloomberg described the current landscape as a "death zone": two weeks ago Moonshot AI's Kimi K3 delivered comparable performance to top US models at a fraction of the cost; now Alibaba is going head-to-head with the frontier at maximum scale. US chip export controls have not stopped Chinese AI companies from matching or surpassing American models on benchmarks.

WangDou's Take

A 2.4-trillion-parameter model scoring 10 points above Sol and nearly 20 above Fable 5 on IFBench — that is not statistical noise, that is a canyon. Of course, two of the benchmarks Alibaba is touting are its own (QwenSWEBench and QwenQoderBench), and ranking first on your own exam does not count. What matters are the third-party tests: IFBench, PaperBench, HealthBench — and on those, Qwen3.8-Max is posting genuinely hard numbers. The more interesting story is price: Chinese model APIs are routinely priced at a fraction of their American counterparts. When performance draws level and the price gap is five to one, the chip-ban narrative gets very hard to sell.

Source: Neowin, Apidog, Bloomberg

Comments

Log in to comment
    This briefing was auto-written by WangDou AI Express for reference only; corrections welcome if you spot a factual error.
    指挥舱👽