15B Supernet Released
Key point
A single checkpoint switches among 8 presets, delivering up to 10.7x faster performance.
Details
SuperApriel-15B-Instruct is a 15B parameter token-mixer supernet that offers 1.0x~10.7x decode throughput at 32K sequence by selecting among 8 deployment presets from a single checkpoint.
- Base model: Apriel-1.6-15b-Thinker
- Training stage 1: stochastic distillation from a frozen teacher, using 266B tokens
- Training stage 2: targeted SFT, using 60B tokens
- Architecture: 48 decoder layers, each layer with 4 mixer variants
- Context length: 262K positions (runtime-dependent)
The presets trade off quality and speed—for example, all-attention shows an average accuracy of 74.2 with 100% quality retention, while Reg|Lklhd-10 drops average accuracy to 57.2 in exchange for a 10.69x speedup.
The benchmarks jointly present multiple categories including math, reasoning, and coding, and emphasize Super Apriel's speed/accuracy balance compared to other hybrid models. For serving, vLLM + Fast-LLM plugin is recommended, and it supports both single-preset mode and supernet mode with runtime switching.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.