Reflection AI Unveils Beam, a 501B MoE Model
Key point
Scheduled for release in October under the Apache 2.0 license, this sparse MoE architecture emphasizes inference compute efficiency.
Details
Reflection AI has announced Beam, a sparse Mixture-of-Experts (MoE) model with 501B total parameters and 23B active parameters per token. As a text-only model specialized for coding, reasoning, and agentic tasks, its weights and technical report are scheduled for release under the Apache 2.0 license in October 2026.
Inference Efficiency and Benchmarks
Beam claims to achieve performance comparable to GLM-5.2 with 3–4x less inference compute. According to the released benchmarks, it outperformed Western open models Inkling and Nemotron 3 Ultra on most metrics.
- SWE Bench Pro v1: Beam 65.5 vs Inkling 54.3
- Terminal Bench v2.1: Beam 80.1 vs Inkling 63.8
- HLE: Beam 36.2 vs Inkling 29.7
However, compared to the similarly sized DeepSeek V4.1 Flash (16B active), it underperformed on 7 out of 9 metrics, suggesting the model focuses on compute efficiency rather than absolute performance.
Training Scale and Infrastructure
Pre-training used 23.8T tokens and was completed in less than 4 weeks using 6,144 NVIDIA GB300 NVL72 GPUs. During the reinforcement learning (RL) phase, 10.5K NVIDIA GB300 GPUs were mobilized to generate over 100 million rollouts.
- Context Extension: Mid-training extended the effective context to 1M tokens.
- Asynchronous Training: Applied Asynchronous Policy Gradient to address Policy Staleness, running an average of 110,000 rollouts concurrently.
- Environment Curation: Secured 1 million high-quality RL environments and operated a separate judge model to prevent reward hacking.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.