AI MODEL PROFILE

Qwen3.8-Max

The highest agentic coding scores currently published by any lab. The open weights shipped on 12 August 2026, but only for a text-only base checkpoint under a custom licence, not the full multimodal API model.

Source review Last reviewed August 22, 2026 By Alibaba Qwen

The facts at a glance

ProviderAlibaba Qwen
TypeFrontier multimodal model
Current version or tierQwen3.8-Max, released 3 August 2026
ModalityText, image, and video in, text out
AccessProprietary API; open weights for the base checkpoint shipped 12 August 2026 under a custom licence, text-only
Context or limits1,000,000 token context, 131,000 max output, 2.4 trillion parameter mixture of experts
Free optionYes, through chat.qwen.ai. Rate limits are not published
Paid priceUSD 2.00 per million input and USD 6.00 per million output, with implicit cache input at USD 0.25

Pricing last checked August 22, 2026. Prices change and vary by region. Check the official pricing page ↗

Assessment

What it does well

  • Alibaba reports 86.6 on Terminal-Bench 2.1, 92.6 on GPQA Diamond, and 86.1 on OSWorld-Verified
  • Accepts video as well as image input
  • Very aggressive pricing for a claimed frontier model, at the same rate as Grok 4.5

Where it falls short

  • Benchmark figures come from Alibaba's own release materials with limited independent replication so far
  • The open-weights checkpoint is text-only and cannot take image or video input, unlike the hosted API model this profile otherwise describes
  • The open checkpoint requires thinking mode and cannot run without reasoning enabled
  • Weights are released under a custom qwen3.8-max licence rather than Apache 2.0, which we have not yet fully reviewed

Privacy and data control

Data handling terms for Qwen Studio and the Qwen API were not verified at the time of writing. Treat the free chat interface as unsuitable for confidential material until you have read the applicable terms.

The longer read

Read the benchmarks with the usual caution

Every figure on this profile comes from Alibaba’s own release materials. That is normal for a model four days old, and it is not a criticism, but self reported scores from any lab are a claim rather than a result until someone else reproduces them. We will revisit this profile when independent evaluations appear.

How it performs by outcome

We assess products against a job, not a leaderboard. These are the outcomes this entry is compared under.

Worth comparing against

How this profile was produced

Source review. Every fact here is drawn from the provider documentation linked below and checked on the date shown. We have not yet run this entry through a controlled test, and no scoring is implied.

Read the full review methodology →

Primary sources