AI Briefing
KO

Zyphra unveils ZAYA1-8B

·2026.05.07 04:43

Key point

Zyphra unveiled ZAYA1-8B, trained on AMD MI300, along with Markovian RSA.

1 / 2

Details

Zyphra unveiled ZAYA1-8B. This model is the first MoE model to perform pretraining, midtraining, and SFT all on the AMD Instinct MI300 stack, boosting math, coding, and reasoning performance with under 1 billion active parameters.

  • Architecture: applies Compressed Convolutional Attention(CCA), an MLP-based expert router, and learned residual scaling.
  • Cluster: trained on an IBM-built cluster using 1,024 MI300x nodes and AMD Pensando Pollara interconnect.
  • Post-training: went through a 5-stage pipeline of SFT → reasoning warmup → RLVE-Gym → math/code RL → RLHF/RLAIF.

Markovian RSA was also presented. It combines parallel multi-trace generation with fixed-length chunking to handle long reasoning, and according to the article's figures it scored 89.6 on HMMT'25, surpassing Claude 4.5 Sonnet and GPT-5-High's 88.3, and on APEX-shortlist it outperformed DeepSeek-V3.2 and GPT-OSS-High at the 5.5M tokens/problem setting.

The model weights were released on Hugging Face and are distributed under the Apache-2.0 license. The service is also offered as a serverless endpoint on Zyphra Cloud.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.