Intern-S2-Mobius Released: Enhancing Efficiency by Separating Knowledge and Reasoning
·2026.08.17 21:49
Key point
Proposes the Mobius-v0 architecture, which improves training efficiency and inference speed by separating knowledge storage and reasoning processes.
Details
Mobius-v0 is an architecture composed of a globally shared Memory (FFN) for storing knowledge vectors and multiple Reasoners (Self-Attn) for performing iterative reasoning.
This structure enhances knowledge compression performance and maximizes inference efficiency by separating knowledge and reasoning.
Key Experimental Results:
- 7B Model (Trained from Scratch): Achieved performance similar to the existing 7B Transformer baseline while using only 62.6% of the training data.
- Intern-S2-Mobius (Continual Pre-training based on Qwen3.5-35B): Improved end-to-end inference speed by approximately 4x while maintaining performance comparable to existing models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.