Fourteen models, one page. Every figure here is pulled from the same profile pages linked below — price, context window and licence, nothing invented for this table. Click any column heading to sort; click a row to open the full review.

What is deliberately not here: a single “strength” score. Vendors report capability on different benchmarks — SWE-bench Pro, FrontierSWE, MCP-Atlas, Terminal-Bench — measured by different methods, so collapsing them into one number would manufacture false precision. Benchmark results stay on each model’s own page, attributed to whoever published them. This page compares only what every vendor publishes as a hard, comparable figure: price and context window.

Full comparison

Prices are per million tokens, published rate as of the “checked” date on each profile.

↔ this box scrolls sideways to show every column
Provider Model Input $/M Output $/M Context Licence Best for
Meta Superintelligence LabsMuse SparkNot publishedNot publishedNot disclosedProprietaryEveryday questions inside apps people already have open
Mistral AIMistral Medium 3.5$1.50$7.50256KOpen weightsCoding agents running on a small GPU cluster
Moonshot AIKimi K3Not publishedNot published1.05MOpen weightsResearch teams that need frontier scale weights they can inspect
Alibaba QwenQwen3.8-Max$2.00$6.001MOpen weightsMulti day autonomous coding projects
DeepSeekDeepSeek-V4-Flash$0.44$1.321MOpen weightsTeams that need to run a capable model on their own infrastructure
xAIGrok 4.5$2.00$6.00500KProprietaryTerminal and agentic coding work where token efficiency matters
Google DeepMindGemini 3.1 Pro$2.00$12.001.05MProprietaryAbstract reasoning problems and research inside the Google ecosystem
Google DeepMindGemini 3.6 Flash$1.50$7.501MProprietaryWork that mixes documents, images, audio, and video in one request
AnthropicClaude Fable 5$10.00$50.001MProprietaryLong horizon tasks where the cost of a wrong answer exceeds the cost of the tokens
AnthropicClaude Sonnet 5$2.00$10.001MProprietaryEveryday reasoning and autonomous tool use without a subscription
AnthropicClaude Opus 5$5.00$25.001MProprietaryAgents that have to keep going for hours without losing the thread
OpenAIGPT-5.6 Luna$0.20$1.201.05MProprietaryHigh volume work where cost per token decides the architecture
OpenAIGPT-5.6 Terra$2.00$12.001.05MProprietaryEveryday production work that still needs the full context window
OpenAIGPT-5.6 Sol$5.00$30.001.05MProprietaryHard reasoning, agentic coding, and security research

Input price vs. output price

Both axes are price per million tokens, log scale. We dropped context window as the second axis: eleven of the twelve priced models now sit within a 500K-to-1.05M-token band, so it barely spread the points apart. Output-to-input ratio does more work — it clusters into three bands by vendor tier (6×, 5×, 3×), visible here as three roughly parallel groups. Hover a point for detail; labels that would overlap are nudged clear with a leader line. Unpriced models are listed separately below.

What would this actually cost?

We’ve argued before that the number to compare is cost per completed task, not price per token. This is that argument made concrete: pick a volume of input text, see what running it through each priced model actually costs.