ZAYA1-74B Preview released
·2026.05.08 07:50
Key point
Zyphra has released ZAYA1-74B-Preview, trained on AMD MI300x.
1 / 2
Details
Zyphra has released ZAYA1-74B-Preview. This model is a MoE-based pre-RL reasoning base checkpoint with 4B active / 74B total parameters, and has not undergone RL post-training or instruction/chat tuning. The weights have been released on Hugging Face under Apache 2.0.
- Pretraining was conducted in 2 stages, training on a total of about 15T tokens across a general knowledge-focused stage and a stage strengthening math, coding, and science capabilities.
- Midtraining continued through 3 stages, extending context from 32k → 128k → 256k, with each stage on the scale of about 1T tokens.
- For long-context efficiency, the architecture replaces every other attention layer with 4K SWA (sliding window attention), and uses Zyphra's CCA attention variant.
- Training was performed on AMD MI300x GPUs with the Pensando Pollara interconnect, and the company stated that expert-context-parallel folding and CCA compute efficiency boosted long-context training efficiency.
Zyphra explained that while this checkpoint is not the final RL model, its pass@4 results are competitive. The company is currently conducting full-scale RL on AMD, and plans to release the final ZAYA1-74B within a few weeks.