On 20 August 2026 a model called Ox Alpha appeared on OpenRouter under the provider name Stealth, with a million-token context window, no price and no stated maker. We covered it at the time and said plainly that the identity was unconfirmed. On 26 August 2026 Z.ai confirmed it. Ox Alpha was a preview of GLM-5.3-Flash, which the company released the same day.
OpenRouter’s own catalogue now carries the answer on both pages. The stealth listing states that the model “was revealed to be ZAI GLM-5.3-Flash”, and the production entry at z-ai/glm-5.3-flash says the same thing in reverse. That is the crediting party’s own record rather than an inference from tokeniser fingerprints, which is the standard this publication applies to attribution claims, and it settles a question we left open six days ago.
The model is a mixture of experts with 320 billion total parameters and 18 billion active, per Z.ai’s model card, and the company describes it as the first natively multimodal model in the GLM-5 series. It accepts text and images; Z.ai’s developer documentation also lists video under visual understanding, though the model card’s own examples cover text and images only. The weights were published to Hugging Face on the day of the announcement.
Why it matters
A capable multimodal model with genuinely permissive terms changes what a buyer can build without first paying for a legal review. MIT imposes no field-of-use restriction, no acceptable-use annexe and no revenue threshold on commercial deployment, and that is not true of the other large open-weight releases of the past month. For teams that had ruled out Chinese open weights on licence grounds specifically, this particular release removes that objection.
The licence is the story, and we read the file
This publication’s standing position is that open weights and open source are different things, and that the difference lives in the licence file rather than in the announcement. So we opened it. The LICENSE file in the zai-org/GLM-5.3-Flash repository is the canonical MIT text, carrying a 2026 copyright held by Z.AI Co., Ltd, with no appended acceptable-use policy, no named-entity exclusion, no downstream naming obligation and no additional conditions of any kind.
That is a materially different position from the recent comparators. Alibaba shipped the Qwen3.8-Max weights under a custom licence, and Kimi K3 arrived under custom terms that require review before commercial use. Those are open-weight releases. This one is open source in the ordinary meaning of the phrase, and the distinction is not pedantry, because it determines whether anyone has to read a document before the model goes into a product.
One caveat that applies to every release of this kind: an MIT grant covers the weights as distributed. It says nothing about the training data behind them, and it is not a warranty against third-party claims. It removes the licence question, not every legal question.
The million-token figure means three different things
Z.ai’s developer documentation advertises support for a one-million-token context window. The Hugging Face model card describes evaluation at a maximum context length of 300,000 tokens. OpenRouter’s production listing routes the model at 1,310,720 tokens. All three numbers are published by parties in a position to know, and all three are accurate, but they are not describing the same quantity.
The routing figure is a ceiling the serving infrastructure will accept. The advertised figure is a round number for the front of the announcement. The evaluation figure is the only one of the three attached to any published measurement of whether the model still behaves at that length. Readers who have followed our argument that a million tokens is not a million tokens will recognise the shape of this. We flag it without treating it as bad faith, because vendors routinely publish all three numbers and label none of them, and Z.ai is no worse than its peers here.
The price on the page this week is not the price
Z.ai’s published list pricing for GLM-5.3-Flash is 0.15 dollars per million input tokens, 0.03 dollars per million cached input tokens and 0.50 dollars per million output tokens. The rates displayed this week are half of each of those figures, at 0.075, 0.015 and 0.25 dollars, because a fifty percent promotion is running. Z.ai’s own pricing page states that the promotion ends at 24:00 on 9 September 2026, Singapore time.
Anyone building a unit-economics model on the numbers visible today should model the doubling that follows. For scale, Z.ai lists the text-only GLM-5.3 at 1.4 dollars per million input tokens and 4.4 dollars per million output, roughly nine times the Flash list input rate and nearly nine times the output rate. Introductory pricing has gone both ways this month across the industry, so the expiry date is worth a calendar entry rather than an assumption in either direction.
The chips are the part worth watching
Z.ai’s documentation states that the company served GLM-5.3-Flash on a large-scale cluster of Chinese AI chips, and claims it reached hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. That is a vendor claim about the vendor’s own infrastructure and we report it as one. There is no independent measurement of the cost comparison, and the comparison itself has no published methodology attached.
The South China Morning Post reported that the stealth preview ran on a cluster of one hundred thousand domestically produced chips and processed sixty-two trillion tokens before the formal release, and that Zhipu’s shares closed more than twelve percent higher at 1,160 Hong Kong dollars. Those are figures from named reporting rather than from the company’s own filings, and we label them accordingly.
Two distinctions matter here and both are easy to lose in retelling. The claim concerns inference, not training. And it concerns the serving of this specific model during this specific preview, not the whole GLM line and not the company’s permanent infrastructure. Several aggregator write-ups this week have already inflated the token figure and blurred the training-versus-inference line; the underlying claim is narrower than the summaries of it.
If it holds in that narrow form, it is still the most consequential thing in the release. A month of Chinese frontier-model news has been about weights and about prices. Serving a million-token multimodal model to a large public audience on non-NVIDIA silicon, at a cost the vendor says is competitive, is a claim about whether export controls constrain deployment at all rather than merely training.
What this means for buyers
If you were blocked on licence terms, you are unblocked. Verbatim MIT weights with no annexe mean this model can go into a commercial product without a bespoke legal review, which is the single largest practical difference between this release and the other large open-weight launches of August 2026. That also means you can self-host it and remove the serving-location question entirely, which is the cleanest answer available if provenance matters to your organisation.
Do not build a budget on this week’s rates. Use the list figures of 0.15 and 0.50 dollars per million tokens, treat the current numbers as a discount expiring on 9 September 2026, and re-check on 10 September. If you are evaluating the context window, test at the length you actually intend to use rather than the advertised ceiling, because 300,000 tokens is the only length Z.ai has published an evaluation against. And if you use the hosted API rather than the weights, the serving infrastructure the vendor describes is in China, which is a data-residency question your procurement team will ask before your engineers do.
What would change our reading
Three things. First, an independent benchmark of long-context behaviour above 300,000 tokens: if quality falls away sharply there, the million-token figure is a routing limit and we will describe it as one. Second, any change to the repository licence. MIT cannot be revoked for weights already distributed, but different terms on a future checkpoint would tell us the permissive posture was a launch tactic rather than a position.
Third, and most important, independent confirmation or contradiction of the domestic-silicon serving claim. That claim currently rests on Z.ai describing its own infrastructure, corroborated by one named report, and it is the load-bearing element of the strategic reading above. A third-party measurement of throughput or per-token cost on that cluster would settle it in either direction.
We are also watching whether the promotion is extended past 9 September 2026, allowed to lapse, or made permanent. Which of the three a lab chooses has become a reasonable tell for how confident it is in its own unit economics.
Sources
- Z.ai, developer documentation for GLM-5.3-Flash stating the one-million-token context window, input modalities and the Chinese-chip serving claim — docs.z.ai
- Z.ai, published API pricing showing list rates, the fifty percent promotion and its 9 September 2026 expiry — docs.z.ai
- Hugging Face, the zai-org/GLM-5.3-Flash model card giving parameter counts and the 300,000-token evaluation context length — huggingface.co
- Hugging Face, the repository LICENSE file, read directly and found to be verbatim MIT — huggingface.co
- OpenRouter, production listing for z-ai/glm-5.3-flash carrying the Ox Alpha reveal, routed context length and current promotional rates — openrouter.ai
- OpenRouter, the original stealth/ox-alpha listing, which now states the model was revealed to be ZAI GLM-5.3-Flash — openrouter.ai
- South China Morning Post, named reporting on the chip cluster size, token volume before release and Zhipu’s share move — scmp.com