Skip to main content
株式会社オブライト
AI2026-08-045 min read

Qwen3.8 Max: 2.4T MoE, $2/M Tokens, Open Weights (2026)

Qwen3.8 Max: Alibaba's 2.4T sparse MoE, ~95B active/token. SWE-bench 87.3%. API from $2.00/$6.00/M tokens (in/out). Open weights promised. Updated Aug. 2026.


What Is Qwen3.8 Max — The Short Answer

Qwen3.8 Max is the flagship model that Alibaba's Qwen team announced on August 3, 2026. It is a sparse mixture-of-experts (MoE) model with 2.4T total parameters, activating roughly 95B parameters per token (an activation ratio of about 4%). It supports up to 1 million tokens of context and up to 128k tokens of output. API pricing starts at $2.00 per million input tokens and $6.00 per million output tokens, with cached input priced at $0.25. Notably, Alibaba has announced that Qwen3.8 Max will be the first Max-class model to have its weights released, and the smaller Qwen3.8-27B will be released as open weights at the same time.

Core Specs

ItemValue
ProviderAlibaba's Qwen team
AnnouncedAugust 3, 2026
ArchitectureSparse MoE (mixture of experts)
Total parameters2.4T
Active parameters per token~95B (activation ratio ~4%)
Context lengthUp to 1,000,000 tokens
Max output128,000 tokens
Weight releaseNot yet released (announced, see below)

Architecture — Only a Sliver of 2.4T Is Active

Qwen3.8 Max uses a sparse mixture-of-experts design: of its 2.4T total parameters, only about 95B — roughly 4% — are actually used for computation on any given token. This design appears intended to balance the representational capacity of a large parameter count against the compute cost of inference. It should be noted, however, that even the 95B active parameter figure is far too large for typical local, personal-hardware inference, and running the full 2.4T model would require substantially larger infrastructure still.

Benchmarks — Where It Lands in Third-Party Evaluations

BenchmarkQwen3.8 Max resultComparison
Frontend Code Arena (overall)#4, Elo 1,668Behind Claude Opus 5 and Kimi K3
Vision Arena#2, Elo 1,305
Vals Index66.1Matches Claude Opus 4.7's score at roughly 1/2.3 the cost
SWE-bench (resolve rate)87.3%Above GPT-5.5's 82.6%

All of these figures come from third-party evaluations. The SWE-bench resolve rate of 87.3% is reported to exceed GPT-5.5's 82.6%. On the Vals Index, Qwen3.8 Max is reported to match Claude Opus 4.7's score at roughly 1/2.3 the cost, positioning it as a cost-performance play.

API Pricing — From $2.00 Input, $6.00 Output per Million Tokens

ItemPrice (per million tokens)
Input$2.00
Output$6.00
Cached input$0.25

This pricing reflects the figures at announcement; whether there is tiered pricing by context length, or regional price differences, has not been confirmed as of this writing. Some other Qwen-series models do use tiered pricing that rises with context length, so it is worth checking the official documentation before committing to heavy usage. For a look at how tiered pricing works on a related Qwen model, see Qwen3.7 Flash's tiered pricing breakdown.

Open Weights Coming — Qwen3.8-27B to Ship Alongside

Qwen3.8 Max is notable as the first Max-class model from Alibaba to have an open-weight release announced, with the release expected "next week." Max-class models have historically been closed and API-only, so this marks a shift in approach. At the same time, the smaller Qwen3.8-27B is also planned for open-weight release. However, its detailed specs — whether it is a dense or MoE architecture, and VRAM requirements at various quantization levels — remain unconfirmed as of this writing, pending an official spec announcement.

- Detailed specs for the 27B model (dense vs. MoE, VRAM by quantization) are unconfirmed
- Local execution of Qwen3.8 Max itself (2.4T) is effectively impractical on personal hardware
- The exact weight-release date is unconfirmed (only "next week" has been stated)
- License terms are unconfirmed

Intended Use Cases — Long-Horizon, Autonomous Agent Work

Alibaba's announcement emphasizes long-horizon, autonomous agentic use cases for Qwen3.8 Max. The specific examples cited include:

- Autonomous coding work sustained over 10+ days
- Chip design optimization processes spanning 500+ turns
- Year-long (365-day) e-commerce strategy planning and execution

How It Compares — Choosing Between Qwen Models

Qwen3.8 Max leads with strong overall coding performance and a pitch for sustained autonomy in agentic work, with its 87.3% SWE-bench resolve rate positioned above GPT-5.5. For vision-heavy or low-cost batch-processing needs, a lower tier within the same family such as Qwen3.7 Flash may offer a better price point. For agentic coding or local-deployment priorities, Qwen3.6 Plus's agentic coding guide and Qwen 3.6 27B (dense architecture) are also worth comparing. Once weights are released, the newly announced Qwen3.8-27B could become a strong candidate for local deployment, but with its detailed specs still unconfirmed, that judgment should be held for now.

FAQ

Can Qwen3.8 Max be run locally?

Running the full 2.4T Max model locally is not practical on typical personal hardware. The smaller Qwen3.8-27B, announced for open-weight release at the same time, may serve as a lighter alternative, but its detailed specs remain unconfirmed.

When will the weights be released?

Alibaba has announced release 'next week,' but the exact date is unconfirmed as of this writing. License terms are also pending official announcement.

What does the API cost?

Pricing starts at $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.25. Whether there is tiered pricing by context length has not been confirmed.

What does 'active parameters' mean?

In a mixture-of-experts (MoE) model, it refers to the number of parameters actually used in computation for a single token's inference. For Qwen3.8 Max, about 95B of the 2.4T total — roughly 4% — are activated per token.

Are the detailed specs of Qwen3.8-27B known?

As of this writing, details such as whether it uses a dense or MoE architecture, and VRAM requirements at different quantization levels, remain unconfirmed pending an official spec release.

Feel free to contact us

Contact Us