Liquid AI Scales Vocabulary Without Retraining the Tokenizer
Key point
Liquid AI has unveiled a technique that expands the vocabulary of the LFM2.5-8B-A1B model from 65K to 128K without retraining from scratch.
Details
Liquid AI has unveiled a Tokenizer Expansion (in-place) technique. It's a methodology that allows upgrading the tokenizer of an existing pretrained model without retraining, and it was applied to the LFM2.5-8B-A1B model.
Key changes:
- Doubled the vocabulary size from 65K → 128K
- Improved processing quality for languages that the existing tokenizer had been over-segmenting
- Reused existing model weights without retraining from scratch
The technical report is available on arXiv (2607.15232), and the model is available on Hugging Face (LiquidAI/LFM2.5-8B-A1B).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.