NVIDIA Unveils Star Elastic
·2026.05.10 09:48
Key point
NVIDIA has unveiled Star Elastic, a technology that can extract models of various sizes from a single checkpoint.
Details
NVIDIA has announced Star Elastic, a new post-training method applied to Nemotron Nano v3. This technology includes three model sizes—30B, 23B, and 12B—within a single checkpoint, and allows instant extraction of a model at the desired size through Zero-Shot Slicing.
Key Technical Features:
- Learnable Router: Via Gumbel-Softmax, it maps various architectural components—attention heads, Mamba SSM heads, MoE experts, FFN channels—to fit the optimal parameter budget.
- Stage-Specific Model Size Optimization: During the inference process, the 23B model is allocated to the 'Thinking' stage and the 30B model to the final 'Answer' stage, maximizing efficiency.
Key Results:
- Performance Improvement: 16% higher accuracy and 1.9x reduced latency compared to standard budget-control methods (based on benchmarks such as AIME-2025 and GPQA).
- Cost Savings: Token usage was reduced by 360x compared to training models from scratch.
- Hardware Accessibility: The 12B NVFP4 variant can run even on an RTX 5080, and achieved a speed of 7,426 tokens per second on the RTX Pro 6000.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.