ZAYA1-8B, R1-level Math Performance
Key point
Zyphra has released ZAYA1-8B, a 760M active parameter MoE model.
Details
Zyphra has released ZAYA1-8B. With an 8.4B total / 760M active MoE architecture, it is a small model that delivers DeepSeek-R1-level math performance based on Zyphra's own evaluation.
On Zyphra's own benchmarks, it recorded AIME 2026 89.1, HMMT 71.6, and LiveCodeBench 65.8. Base score and RSA-boosted score were presented separately, and all figures are Zyphra's own evaluations.
- Markovian RSA doesn't stack a long reasoning chain all at once; instead it splits it into chunks, passing only the tail of each step as the seed for the next inference.
- This approach was designed to limit the context window while still leveraging more test-time compute.
- Zyphra trained the model and the inference method together, and stated that attaching the same RSA to Qwen3-4B produces a smaller performance gain.
Training was conducted entirely on a 1,024-node cluster based on AMD Instinct MI300X. Pretraining, midtraining, and supervised fine-tuning all used only the AMD stack.
However, general-purpose agent performance is weak. Tool calling was weak with BFCL-V4 39.22 and TAU2 43.12, and IFBench 52.56, EQBench 72.95, and Creative Writing 62.97 are also not strengths.
Weights have been released on Hugging Face under Apache 2.0, and it's also available on Zyphra Cloud. Local inference requires Zyphra's fork of vLLM.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.