
OX Alpha: The Mystery AI Model That Beats GPT-5.6 in Code — Probably Chinese
Appearing anonymously on OpenRouter on August 20, OX Alpha outperforms GPT-5.6 Sol and Claude Fable 5 on code benchmarks — 80% on DeepSWE. Free until August 27, 1M context tokens, multimodal. Investigation into a fifth stealth model that clearly points to Z.AI.
On August 20, 2026, an AI model named “stealth/ox-alpha” appeared without any branding on OpenRouter and OpenCode. The fifth stealth model released on these platforms in six months. But this one immediately made waves: it crushes GPT-5.6 Sol and Claude Fable 5 on several code benchmarks. And it is free until August 27.
Announced Specs
- Context: 1,048,576 tokens (1M)
- Multimodal input: text, image, video
- Function calling supported
- Input AND output tokens free during the public preview
The Buzz-Worthy Performance
Developer Ben Davis runs OX Alpha through DeepSWE, a code agent benchmark, and publishes his results: 80% for OX Alpha versus 65% for Claude Fable 5 and 52% for GPT-5.6 Sol. In less than 24 hours, editors like Cursor and Zed integrate it in A/B testing for their code assistant.
An unknown model, in free public preview, dominating the market's paid leaders in code. Such an event hasn't happened since DeepSeek V3.
The Important Caveat: No Official Benchmark
The 80% figure comes from a single, small-sample test published by an independent researcher. No official benchmark is available to date. Several other community evaluations follow — results vary between 65 and 85% depending on configurations and seeds. These figures should be treated as preliminary signals, not as a definitive ranking.
From experience, real gaps are seen in edge cases (complex bugs, massive refactorings, very long contexts) — not in traditional academic benchmarks.
Where Does OX Alpha Come From?
The question fascinates the ecosystem. Three hypotheses circulate:
- Z.AI (Zhipu AI, China) — Ben Davis claims to be “99% certain” it is a new variant of the GLM-5.3 series. Stylistic signatures, response format, benchmarks very close to GLM-5.2 Turbo released on August 17.
- Anthropic or xAI testing Claude Opus 5.5 or Grok 5 anonymously in public.
- New independent Chinese lab (Baichuan, MiniMax, Moonshot) attempting a viral launch.
The Z.AI theory is the most documented. The pattern is consistent with the Chinese strategy: release in stealth, let the community validate performance, then officially brand it once the buzz is confirmed.
Why It Matters
Three interpretations intersect:
- American dominance in code is ending. Benchmarks held by OpenAI and Anthropic since 2023 are being eroded by Chinese players who iterate faster and publish in open-weights.
- Zero pricing temporarily kills the market. Cursor, Windsurf, Continue.dev must revise their model mix when a superior tool is free. The margin for intermediaries shrinks further.
- Stealth releases become standard practice. This is the fifth in six months. Major labs test their models in real conditions before the official announcement — a practice imported from the video game world.
What to Watch
The free preview ends on August 27. Two scenarios:
- OX Alpha becomes the official GLM-5.3 of Z.AI in the following days, with a pricing grid and API confirmation — this is the likely path.
- The model disappears without explanation, which would support a more discreet origin (research lab testing in public conditions).
Also to verify: the independent third-party benchmarks (LiveCodeBench, SWE-Bench Verified, MMLU) that will drop in the coming days and confirm — or deflate — the announced performance.
One thing is certain: between Kimi K3, Qwen 3.8 Max, GLM-5.2, DeepSeek V4, and now OX Alpha, the layer of Chinese open-weights code models is becoming the most dynamic in the world in 2026. French developers would be well advised to test it.