Zyphra unveils ZAYA1-8B
Key point
Zyphra unveiled ZAYA1-8B, trained on AMD MI300, along with Markovian RSA.
Details
Zyphra unveiled ZAYA1-8B. This model is the first MoE model to perform pretraining, midtraining, and SFT all on the AMD Instinct MI300 stack, boosting math, coding, and reasoning performance with under 1 billion active parameters.
- Architecture: applies
Compressed Convolutional Attention(CCA), an MLP-based expert router, and learned residual scaling. - Cluster: trained on an IBM-built cluster using 1,024 MI300x nodes and AMD Pensando Pollara interconnect.
- Post-training: went through a 5-stage pipeline of SFT → reasoning warmup → RLVE-Gym → math/code RL → RLHF/RLAIF.
Markovian RSA was also presented. It combines parallel multi-trace generation with fixed-length chunking to handle long reasoning, and according to the article's figures it scored 89.6 on HMMT'25, surpassing Claude 4.5 Sonnet and GPT-5-High's 88.3, and on APEX-shortlist it outperformed DeepSeek-V3.2 and GPT-OSS-High at the 5.5M tokens/problem setting.
The model weights were released on Hugging Face and are distributed under the Apache-2.0 license. The service is also offered as a serverless endpoint on Zyphra Cloud.