Fermion Research Unveils Neutrino-1 Model Family
Key point
The Neutrino-1 model family has been released, applying a ternary-style format from the training stage to maximize efficiency.
Details
Fermion Research has unveiled a family of 3 Neutrino-1 models that directly apply a Ternary-style format starting from the training stage. Rather than the conventional post-training quantization approach, these models are designed via Native-format training, where all forward passes occur within the low-resolution format.
Key Model Lineup:
- Neutrino-1 8B: An 8B-class flagship model (MMLU 72.1)
- Neutrino-1 0.6B: A draft model and small model
- Neutrino-1 0.6B-Chat: A chat-specialized model
Core Technology and Features:
- Storage and Bandwidth Savings: The 8B model is highly efficient, with a 3.88GB disk footprint and a 2.56GB download size, addressing the bandwidth bottleneck that arises during single-stream decoding.
- Unified Container Architecture: All three models share the same tokenizer, container layout, and inference engine, allowing the 0.6B model to be used immediately as a Draft model for the 8B model.
- GQA (Grouped-Query Attention) Applied: An efficient attention mechanism has been adopted.
- Open Source: All models are licensed under Apache License 2.0 and are readily available on Hugging Face.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.