SupraLabs Releases Supra-50M at 50M Scale
Key point
SupraLabs has released Supra-50M, a small language model with 50M parameters trained on 20 billion tokens.
Details
SupraLabs has unveiled Supra-50M, a 50M-parameter Causal Language Model based on a Llama-style architecture. The model is available in two versions: Base and Instruct.
Key Features and Training Information:
- Training Data: Uses 20 billion (20B) tokens of high-quality educational web text from
fineweb-edu - Architecture: Llama-style Decoder-only Transformer (with GQA applied)
- Tokenizer: Custom Byte-Level BPE trained by sampling 500,000 documents
Benchmark Performance: Despite its very small model size, Supra-50M demonstrates performance competitive with larger-scale models.
- BLiMP (linguistics): 76.3% (surpassing GPT-2 124M model's 63.0%)
- SciQ (science): 77.2%
- ARC-Easy (knowledge): 52.2%
Starting with this model, SupraLabs plans to sequentially release Supra-124M and Supra-350M models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.