Qwen3.8
Alibaba
Alibaba’s Qwen3.8-Max, previewed 19 July and launched 3 August 2026 as the largest model in the Qwen family: a mixture-of-experts design with 2.4 trillion total parameters and 95 billion active per query, across 512 experts. Alibaba published the weights on 12 August as Qwen/Qwen3.8-2.4T-A95B under a licence named "qwen3.8-max" - not Apache 2.0 - making it the largest downloadable model on our roster. It is text-only despite widespread reporting to the contrary: the multimodal member of the family is the separate Qwen3.8-27B. Context is 262,144 tokens natively, extensible to about 1.01M, and thinking mode is always on rather than optional.
Pricing & specs verifiedSource: vendor documentation · checked Sep 26, 2026LMArena: 1479 (rank 25) · leaderboard · retrieved Sep 26, 2026
Pricing
$2 / $6 per 1M tokens
Context
262k native (extensible to ~1.01M)
LMArena Elo
1479
Overall rank
#25
Key Features
Multimodal Input Video Understanding Agentic Workflows Long Context 2.4T MoE
Full Specifications
$2 / $6 per 1M tokens | |
|---|---|
| Identity | |
| Developer | Alibaba |
| Released | Aug 2026 |
| Status | GA |
| Licence | Open weights under a licence named "qwen3.8-max" - NOT Apache 2.0 (the Qwen3.8-27B sibling is Apache 2.0, which is a common source of confusion) |
| Self-hostable | Yes - Qwen/Qwen3.8-2.4T-A95B and an FP8 checkpoint published 12 Aug 2026 |
| Cost | |
| Blended $/1M tokensinput × 0.75 + output × 0.25 | $3.000 / 1M tokens |
| Input price | $2.000 / 1M tokens |
| Output price | $6.000 / 1M tokens |
| Cached input | Not documented |
| Batch discount | Not documented |
| Free tier | Not documented |
| Capacity | |
| Context window | 262k native (extensible to ~1.01M) |
| Max output | 128k tokens |
| Long-context surcharge | Not documented |
| Capability | |
| Vision in | No - text-only. Alibaba’s model card states "Multimodal inputs are not supported". The multimodal sibling is Qwen3.8-27B (Apache 2.0), which is tagged Image-Text-to-Text and does accept images and video. |
| Audio in | No |
| Function calling | Yes |
| Structured output | Yes |
| Extended reasoning | Yes - thinking mode is required for all interactions, not optional |
| Web search | Not documented |
| Code execution | Not documented |
| Access | |
| API | Yes |
| Chat app | Yes |
| Cloud marketplaces | Alibaba Cloud Model Studio |
| Fine-tuning | Yes - open weights permit it |
| Measured quality | |
| LMArena Elo | 1479 (max effort, checked Sep 2026) |
| LMArena rank | Rank 25 |
| Elo per dollarLMArena Elo ÷ blended $/1M tokens | 493 |