AI Briefing
KO

Model Size Scaling Outlook 2023-2031

·2026.06.23 09:00

Key point

Based on HBM bandwidth and pipeline architecture, this forecasts how model parameter sizes will change through 2031.

Details

Token generation speed is limited by the read speed of HBM (High Bandwidth Memory), that is, the speed at which model weights and the KV-cache are read. As models grow larger, HBM read time increases, which directly constrains both the number of pipeline stages and the total number of parameters in the model.

The changes in read time according to HBM technology advancement are as follows.

  • H100: 20ms
  • H200: 30ms
  • GB200/GB300: 24ms / 36ms
  • Rubin / Rubin Ultra: 13ms
  • Feynman (2028-2029): 16ms / 14ms

Based on this hardware performance, the forecasted model size starts at 10T in 2026, reaches 240T in 2028, and is expected to reach 1.4 quadrillion parameters by 2031. Notably, from 2027 onward, due to training data shortages, the 2031 model is expected to need to be 4 times larger than it would be under an unlimited data assumption.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.