What’s converging across AI, what’s pulling apart, and what’s just starting to show up. Every claim here traces to the same model and app data behind our other pages, or to a story we’ve already reported.
Context windows have stopped being a differentiator
11 of 12 priced models in our directory now sit within a single band — 500,000 to 1.05 million tokens. Two years ago, context window was the headline spec labs competed on. It no longer separates anyone meaningfully; nearly every frontier lab converged on roughly the same number within the same year.
What that means practically: if you’re choosing a model on context window alone, you’re choosing on a spec that barely varies anymore. The real differentiation moved elsewhere — see our explainer on what a million-token context actually costs you, which was already making this point before the convergence was this complete.
Pricing strategy is splitting into distinct tiers, not converging
While context windows converged, pricing did the opposite. Plotting input price against output price across our fourteen tracked models shows three distinct bands by output-to-input ratio: 6× for OpenAI’s flagship tier and Gemini 3.1 Pro, 5× for Anthropic, Mistral, and Gemini 3.6 Flash, and 3× for the cheaper open-weight tier — Qwen, DeepSeek, Grok.
That’s a real strategic split, not noise: labs pricing output at 6× input are betting buyers will pay a premium for finished work regardless of how many tokens it took to get there; labs at 3× are competing on raw throughput cost. See the full breakdown on our model comparison page.
The same divergence shows up in licensing. 4 of 14 models in our directory ship with genuinely open weights; the other ten are proprietary. That split has held steady even as capability has converged — being open or closed has become a durable strategic choice for a lab, not a temporary phase every model passes through on the way to being closed.
Worth reading alongside this: our piece on why open weights and open source are not the same claim, and how the centre of gravity in open weights shifted away from Meta this year.
Stealth launches are becoming a standard distribution channel
Anonymous or pseudonymous model releases on routing platforms — appearing with no confirmed maker, often free during a preview window — are no longer a one-off curiosity. We covered one directly in August: a million-token endpoint with zero pricing and no attributable operator, reported to be at least the fifth such release in recent months.
The pattern is commercially rational — free real-world traffic and preference data before a lab commits to a name — and it creates a real gap for buyers: an endpoint with no accountable operator is one you can query but shouldn’t route production traffic through. Worth watching whether platforms start requiring disclosed operators as this becomes more common.
Capability claims are outrunning verification infrastructure
Across several stories this month, the same failure pattern recurred: a genuinely impressive number circulates widely before anyone checks what it actually measures. A refusal-rate metric got reported as a capability score. A May funding projection got reported as an August result. A benchmark figure with no published harness got repeated by outlet after outlet.
None of these were fabrications — they were real numbers, attached to the wrong meaning, moving faster than anyone checked. This is the specific gap our watchlist exists to close: dated, falsifiable questions, resolved in public, including the times we were wrong.