AI Briefing
KOSign in

Ai2 Releases Olmo-core 3: Open Infrastructure Benchmarked for Trillion-Parameter MoE Training

·2026.10.02 00:01

Key point

Ai2 releases Olmo-core 3, an open-source training framework designed for large MoE models, achieving 2.7x throughput improvement in preliminary tests and benchmarked at over 1 trillion parameters.

1 / 4

Details

Ai2 has released Olmo-core 3, a significant upgrade to its open-source training framework designed to scale Mixture-of-Experts (MoE) models into the trillion-parameter range while maintaining computational efficiency. This infrastructure serves as the foundation for the next generation of Olmo models, which will adopt an MoE architecture.

Key Performance Metrics

The framework introduces a redesigned training stack that switches from fully sharded data parallelism (FSDP) to distributed data parallelism (DDP), keeping experts resident on GPUs to avoid repeated weight gathering. In preliminary tests on eight NVIDIA B300 GPUs:

  • A 47-billion-parameter MoE processed 52,000 tokens per second per GPU, compared to 19,400 tokens per second with the previous implementation, representing a 2.7x throughput increase.
  • Scaling the expert pool from 8 to 128 while selecting only four experts per token kept active parameters at ~3.2B and increased total capacity to 47B, with training throughput dropping by less than 5%.

Technical Optimizations

Olmo-core 3 integrates several techniques to distribute large MoEs across GPU clusters:

  • Expert and Pipeline Parallelism: Spreads experts and model layers across GPUs to reduce memory requirements per device.
  • Distributed Optimizer: Distributes optimizer state across GPUs rather than storing full copies on each.
  • MXFP8 Support: Using this lower-precision format on four NVIDIA B300 GPUs increased training throughput by ~21% compared to BF16, while reducing peak active memory from 103 GiB to 95 GiB.
  • Routing Efficiency: Features like rowwise expert parallelism and GPU-resident routing minimize data rearrangement and CPU waiting times.

Scaling Limits and Findings

The system has been benchmarked at over one trillion total parameters, with a specific test reaching 1.2 trillion parameters (58.36 billion active per token) across 512 GPUs, achieving 858 TFLOP/s/GPU. A short-capacity test using DeepEP v2 demonstrated scalability up to 2.38 trillion parameters, though this was not a full training run.

The accompanying technical report also highlights critical findings, such as "token gerrymandering" (where routing balance scores improve while actual workload balance degrades) and the observation that overlapping communication and computation does not always yield faster end-to-end execution.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.