Snaxbyte-sized news, on the go
Nvidia enlisted six Wall Street firms to fund AI data centres it will not put on its own books. The number is a 500 billion dollar target, not money in the ground. Andreessen Horowitz raised 1.1 billion dollars to fund AI hardware. The firm that said software eats the world is now buying the plumbing. OpenAI is cutting off Cursor’s access to its models. Model supply is now a lever labs can pull against rivals they distrust. OpenAI retired the o3 model family from ChatGPT on 26 August 2026. The move closes a ninety-day sunset window and pushes remaining reasoning-model users toward GPT-5.6 or the pricier o3-pro tier. OpenAI’s one-year deprecation clock on the Assistants API reaches zero today. The company’s own documentation shows no sign of a reprieve, even as independent confirmation that calls are now failing remains thin. Nvidia is reportedly raising AI server prices more than fifteen percent. Memory, not the GPU, is now the cost that moves the bill. Nvidia enlisted six Wall Street firms to fund AI data centres it will not put on its own books. The number is a 500 billion dollar target, not money in the ground. Andreessen Horowitz raised 1.1 billion dollars to fund AI hardware. The firm that said software eats the world is now buying the plumbing. OpenAI is cutting off Cursor’s access to its models. Model supply is now a lever labs can pull against rivals they distrust. OpenAI retired the o3 model family from ChatGPT on 26 August 2026. The move closes a ninety-day sunset window and pushes remaining reasoning-model users toward GPT-5.6 or the pricier o3-pro tier. OpenAI’s one-year deprecation clock on the Assistants API reaches zero today. The company’s own documentation shows no sign of a reprieve, even as independent confirmation that calls are now failing remains thin. Nvidia is reportedly raising AI server prices more than fifteen percent. Memory, not the GPU, is now the cost that moves the bill.
Model prices /1M out
Qwen3.8-Max $6.00
Grok 4.5 $6.00
Mistral Medium 3.5 $7.50
Gemini 3.6 Flash $7.50
Claude Sonnet 5 $10.00
GPT-5.6 Terra $12.00
Claude Opus 5 $25.00
Claude Fable 5 $50.00
MODEL INTELLIGENCE

AI models, explained by what they are actually good at

A working directory of the models that matter right now. Every entry names its provider, states one clear verdict, links the primary source, and shows the date we last checked it. Models change quickly, so treat the reviewed date as part of the fact.

GPT-5.6 Sol

OpenAI

Source review

OpenAI's top tier model, and the one to reach for when the problem is genuinely hard rather than merely long.

Best for
Hard reasoning, agentic coding, and security research
Modality
Text and image in, text out
Access
Proprietary
Context
1,050,000 token context, 128,000 max output, knowledge cutoff 16 February 2026
Free route
No free route. Excluded from the API free tier and from ChatGPT Free at launch

GPT-5.6 Terra

OpenAI

Source review

The sensible default in the GPT-5.6 family: the same very large context as Sol at a fraction of the price.

Best for
Everyday production work that still needs the full context window
Modality
Text and image in, text out
Access
Proprietary
Context
1,050,000 token context, 128,000 max output
Free route
Reached by ChatGPT Free and Go users through ChatGPT Work and Codex at launch

GPT-5.6 Luna

OpenAI

Source review

The cheapest way to get a million token context window from a major lab, after an 80 percent price cut in July 2026.

Best for
High volume work where cost per token decides the architecture
Modality
Text and image in, text out
Access
Proprietary
Context
1,050,000 token context, 128,000 max output
Free route
Yes. Became the default model for ChatGPT Free and Go users in the week of 6 August 2026

Claude Opus 5

Anthropic

Source review

Anthropic's strongest model for work that runs long, where holding context and recovering from mistakes matters more than a single clever answer.

Best for
Agents that have to keep going for hours without losing the thread
Modality
Text and vision, with computer use
Access
Proprietary
Context
1,000,000 token context, 128,000 max output, knowledge cutoff May 2026
Free route
No. Paid Claude plans, Claude Code, the API, AWS Bedrock, Google Cloud, and Microsoft Foundry

Claude Sonnet 5

Anthropic

Source review

The most capable model most people can reach without paying anything, and the default on the free Claude plan.

Best for
Everyday reasoning and autonomous tool use without a subscription
Modality
Text and vision
Access
Proprietary
Context
1,000,000 token context, 128,000 max output, knowledge cutoff January 2026
Free route
Yes. The default model on the free Claude plan

Claude Fable 5

Anthropic

Source review

Anthropic's most capable general model on its own benchmarks, priced accordingly, and worth it only when the task genuinely warrants it.

Best for
Long horizon tasks where the cost of a wrong answer exceeds the cost of the tokens
Modality
Text and vision
Access
Proprietary
Context
1,000,000 token context, 128,000 max output, knowledge cutoff January 2026, adaptive thinking always on
Free route
No ongoing free tier

Gemini 3.6 Flash

Google DeepMind

Source review

The broadest input coverage of any current model, and the practical choice when your source material is not just text.

Best for
Work that mixes documents, images, audio, and video in one request
Modality
Text, image, video, audio, and PDF in, text out
Access
Proprietary
Context
1,000,000 token input, 64,000 max output
Free route
Yes. Free tier on the Gemini API, plus the Gemini app and Google AI Studio

Gemini 3.1 Pro

Google DeepMind

Source review

Google's reasoning flagship, still shipping under a preview model ID nearly six months after release.

Best for
Abstract reasoning problems and research inside the Google ecosystem
Modality
Multimodal in, text out
Access
Proprietary
Context
1,048,576 token context, 65,536 max output
Free route
No free API tier. Consumer access through the Gemini app and NotebookLM

Grok 4.5

xAI

Source review

Strong published results on agentic coding benchmarks using far fewer tokens than rivals, in a text only package.

Best for
Terminal and agentic coding work where token efficiency matters
Modality
Text in, text out
Access
Proprietary
Context
500,000 token context, knowledge cutoff 1 February 2026
Free route
Yes, for a limited time in Grok Build and Cursor, plus grok.com, X, iOS, and Android

DeepSeek-V4-Flash

DeepSeek

Source review

A genuinely permissive MIT licence on a model with a million token context, at roughly a thirtieth of the price of the Western flagships.

Best for
Teams that need to run a capable model on their own infrastructure
Modality
Text
Access
Open weights under the MIT licence
Context
1,000,000 token context, 384,000 max output, 304 billion total parameters in a mixture of experts design
Free route
Yes, by self hosting the MIT licensed weights from Hugging Face

Qwen3.8-Max

Alibaba Qwen

Source review

The highest agentic coding scores currently published by any lab. The open weights shipped on 12 August 2026, but only for a text-only base checkpoint under a custom licence, not the full multimodal API model.

Best for
Multi day autonomous coding projects
Modality
Text, image, and video in, text out
Access
Proprietary API; open weights for the base checkpoint shipped 12 August 2026 under a custom licence, text-only
Context
1,000,000 token context, 131,000 max output, 2.4 trillion parameter mixture of experts
Free route
Yes, through chat.qwen.ai. Rate limits are not published

Kimi K3

Moonshot AI

Source review

The first open 3 trillion parameter class model, released under a custom licence that needs reading before you build on it.

Best for
Research teams that need frontier scale weights they can inspect
Modality
Text, image, and video in, text out
Access
Open weights under the bespoke Kimi K3 Licence, which is not an OSI approved licence
Context
1,048,576 token context, 2.8 trillion total parameters with 104 billion activated across 896 experts
Free route
Yes, by downloading the weights from Hugging Face

Mistral Medium 3.5

Mistral AI

Source review

A dense 128 billion parameter model designed to run on as few as four GPUs, which makes self hosting realistic rather than theoretical.

Best for
Coding agents running on a small GPU cluster
Modality
Text and vision
Access
Open weights under a modified MIT licence
Context
256,000 token context, 128 billion parameters, dense rather than mixture of experts
Free route
Yes, weights on Hugging Face. Also in Le Chat and Mistral Vibe on paid plans

Muse Spark

Meta Superintelligence Labs

Source review

Meta's replacement for the Llama line, distributed free at enormous scale and documented almost not at all.

Best for
Everyday questions inside apps people already have open
Modality
Text, image, and voice
Access
Proprietary
Context
Context window not disclosed
Free route
Yes, through the Meta AI app, meta.ai, WhatsApp, Instagram, Facebook, Messenger, and Meta AI glasses