← Back to articles
OX Alpha: The Mystery AI Model That Beats GPT-5.6 in Code — Probably Chinese

OX Alpha: The Mystery AI Model That Beats GPT-5.6 in Code — Probably Chinese

Appearing anonymously on OpenRouter on August 20, OX Alpha outperforms GPT-5.6 Sol and Claude Fable 5 on code benchmarks — 80% on DeepSWE. Free until August 27, 1M context tokens, multimodal. Investigation into a fifth stealth model that clearly points to Z.AI.

By Rédaction Gennn··2 min read

On August 20, 2026, an AI model named “stealth/ox-alpha” appeared without any branding on OpenRouter and OpenCode. The fifth stealth model released on these platforms in six months. But this one immediately made waves: it crushes GPT-5.6 Sol and Claude Fable 5 on several code benchmarks. And it is free until August 27.

Announced Specs

  • Context: 1,048,576 tokens (1M)
  • Multimodal input: text, image, video
  • Function calling supported
  • Input AND output tokens free during the public preview

The Buzz-Worthy Performance

Developer Ben Davis runs OX Alpha through DeepSWE, a code agent benchmark, and publishes his results: 80% for OX Alpha versus 65% for Claude Fable 5 and 52% for GPT-5.6 Sol. In less than 24 hours, editors like Cursor and Zed integrate it in A/B testing for their code assistant.

An unknown model, in free public preview, dominating the market's paid leaders in code. Such an event hasn't happened since DeepSeek V3.

The Important Caveat: No Official Benchmark

The 80% figure comes from a single, small-sample test published by an independent researcher. No official benchmark is available to date. Several other community evaluations follow — results vary between 65 and 85% depending on configurations and seeds. These figures should be treated as preliminary signals, not as a definitive ranking.

From experience, real gaps are seen in edge cases (complex bugs, massive refactorings, very long contexts) — not in traditional academic benchmarks.

Where Does OX Alpha Come From?

The question fascinates the ecosystem. Three hypotheses circulate:

  • Z.AI (Zhipu AI, China) — Ben Davis claims to be “99% certain” it is a new variant of the GLM-5.3 series. Stylistic signatures, response format, benchmarks very close to GLM-5.2 Turbo released on August 17.
  • Anthropic or xAI testing Claude Opus 5.5 or Grok 5 anonymously in public.
  • New independent Chinese lab (Baichuan, MiniMax, Moonshot) attempting a viral launch.

The Z.AI theory is the most documented. The pattern is consistent with the Chinese strategy: release in stealth, let the community validate performance, then officially brand it once the buzz is confirmed.

Why It Matters

Three interpretations intersect:

  1. American dominance in code is ending. Benchmarks held by OpenAI and Anthropic since 2023 are being eroded by Chinese players who iterate faster and publish in open-weights.
  2. Zero pricing temporarily kills the market. Cursor, Windsurf, Continue.dev must revise their model mix when a superior tool is free. The margin for intermediaries shrinks further.
  3. Stealth releases become standard practice. This is the fifth in six months. Major labs test their models in real conditions before the official announcement — a practice imported from the video game world.

What to Watch

The free preview ends on August 27. Two scenarios:

  • OX Alpha becomes the official GLM-5.3 of Z.AI in the following days, with a pricing grid and API confirmation — this is the likely path.
  • The model disappears without explanation, which would support a more discreet origin (research lab testing in public conditions).

Also to verify: the independent third-party benchmarks (LiveCodeBench, SWE-Bench Verified, MMLU) that will drop in the coming days and confirm — or deflate — the announced performance.

One thing is certain: between Kimi K3, Qwen 3.8 Max, GLM-5.2, DeepSeek V4, and now OX Alpha, the layer of Chinese open-weights code models is becoming the most dynamic in the world in 2026. French developers would be well advised to test it.