AI Briefing
KO

Toto 2.0: Time Series Forecasting Enters the Scaling Era

·2026.05.15 09:00

Key point

Datadog released Toto 2.0, ranging from 4m to 2.5B parameters, demonstrating clear scaling effects.

1 / 2

Details

Datadog has released Toto 2.0, a family of open-weight time series forecasting models ranging from 4m to 2.5B parameters. This version confirms that performance keeps improving as the model gets larger, with no signs of saturation even at 2.5B.

Benchmark performance was also strong. The models ranked at the top across BOOM, GIFT-Eval, and TIME, and on BOOM, every size landed on the Pareto frontier. On GIFT-Eval, the three largest models maintained the lead among foundation models, and on the overall leaderboard, the 2.5B FT and the Toto 2.0 Family and Friends (FnF) ensemble took 1st and 2nd place.

Training data was centered on observability and synthetic data, and no public forecasting data was used during pretraining. Despite this, the models generalized well even to general-purpose benchmarks, and compared to Toto 1.0, the number of parameters needed to achieve the same quality was reduced by about 7x, with inference also becoming much faster.

Evaluation looked at both CRPS and MASE together. CRPS measures the quality of the probabilistic forecast distribution, while MASE measures point-forecast accuracy relative to a seasonal naive baseline, enabling comparisons that are closer to production environments.

Future challenges identified include further narrowing the gap with classical baselines at long horizons, improving data curation, designing evaluations that reflect downstream value, and expanding into multimodality. The model weights and the distributed training infrastructure library dd_unit_scaling have both been released under Apache 2.0, and a technical report covering the training data, architecture, and the u-μP hyperparameter transfer pipeline is expected soon.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.