Qwen3.8-Max
The highest agentic coding scores currently published by any lab. The open weights shipped on 12 August 2026, but only for a text-only base checkpoint under a custom licence, not the full multimodal API model.
The facts at a glance
Pricing last checked August 22, 2026. Prices change and vary by region. Check the official pricing page ↗
Assessment
What it does well
- Alibaba reports 86.6 on Terminal-Bench 2.1, 92.6 on GPQA Diamond, and 86.1 on OSWorld-Verified
- Accepts video as well as image input
- Very aggressive pricing for a claimed frontier model, at the same rate as Grok 4.5
Where it falls short
- Benchmark figures come from Alibaba's own release materials with limited independent replication so far
- The open-weights checkpoint is text-only and cannot take image or video input, unlike the hosted API model this profile otherwise describes
- The open checkpoint requires thinking mode and cannot run without reasoning enabled
- Weights are released under a custom qwen3.8-max licence rather than Apache 2.0, which we have not yet fully reviewed
Privacy and data control
Data handling terms for Qwen Studio and the Qwen API were not verified at the time of writing. Treat the free chat interface as unsuitable for confidential material until you have read the applicable terms.
The longer read
Read the benchmarks with the usual caution
Every figure on this profile comes from Alibaba’s own release materials. That is normal for a model four days old, and it is not a criticism, but self reported scores from any lab are a claim rather than a result until someone else reproduces them. We will revisit this profile when independent evaluations appear.
How it performs by outcome
We assess products against a job, not a leaderboard. These are the outcomes this entry is compared under.
Worth comparing against
How this profile was produced
Source review. Every fact here is drawn from the provider documentation linked below and checked on the date shown. We have not yet run this entry through a controlled test, and no scoring is implied.
Read the full review methodology →