Reflection AI Introduces Beam: A 501B-Parameter Open-Weight MoE Model Optimized for High-Compute RL
Key point
Reflection AI has introduced Beam, a 501B-parameter open-weight MoE model trained on 10.5K NVIDIA GB300 GPUs, claiming high inference efficiency compared to larger models like GLM 5.2 and Kimi K3.
Details
Reflection AI has announced Beam, its first open-weight model, featuring a sparse Mixture-of-Experts (MoE) architecture with 501B total parameters and 23B active parameters. Designed for coding, reasoning, and agentic workloads, Beam claims to be competitive with larger open-weight models while offering superior inference efficiency.
Architecture and Pretraining
Beam was pretrained on 23.8 trillion high-quality tokens from web and proprietary licensed datasets. The model employs an architecture with interleaved local/global attention, fine-grained routed experts, and controlled residual streams to ensure stable optimization. Reflection AI reports that Beam Base matches or exceeds similar-sized open-source base models, with pretraining completed in under four weeks on a cluster of 6,144 NVIDIA GB300 NVL72 GPUs.
Reinforcement Learning and Infrastructure
The model's capabilities are driven by a massive reinforcement learning (RL) investment, described as one of the largest RL runs by an open lab to date. Key infrastructure metrics include:
- Compute: Utilization of 10.5K NVIDIA GB300 GPUs over four weeks.
- Data: Generation of over 100 million rollouts and training on approximately 1.3 billion sandboxes.
- Efficiency: A fully asynchronous execution system where new weights reach the inference fleet in a median of 12 seconds, reducing cross-rack traffic by 75%.
- Stability: The system handled 71 inference incidents without terminating training jobs, recovering capacity in a median of 8 minutes.
Performance and Efficiency Claims
Reflection AI positions Beam as a leader in Western open-weight frontier progress, particularly in inference efficiency. The company claims Beam achieves advanced reasoning scores similar to GLM 5.2 while using 3–4x less inference compute. Compared to 2T+ parameter models like Qwen 3.8-Max, Beam reportedly provides more intelligence per token. While models like Kimi K3 remain ahead on raw capability, Beam's advantage lies in its efficiency at inference time.
Benchmark comparisons provided by Reflection AI show:
- DeepSWE v1.1: Beam 44.4 (vs. GLM 5.2: 44.0, Kimi K3: 68.0)
- SWE Bench Pro v2-Hard: Beam 77.2 (vs. GLM 5.3: 84.3, Kimi K3: 88.2)
- Terminal Bench v2.1: Beam 80.1 (vs. GLM 5.2: 81.0, Kimi K3: 88.3)
- AIME 2026: Beam 97.8 (vs. GLM 5.2: 99.2)
- GPQA Diamond: Beam 90.5 (vs. Kimi K3: 93.5, Qwen 3.8 Max: 92.6)
Availability and Safety
Beam is currently available to select users via a waitlist. Reflection AI plans to release the model weights, technical report, model card, and developer artifacts under the Apache 2.0 license this month. The release will include integration with various open-source libraries and harnesses. Safety alignment was achieved through Multi-Teacher On-Policy Distillation (MOPD) and deliberative alignment techniques, with safety evaluation results to be published in the technical report.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.