Language disappears
Key point
8 languages and math/code converged into the same internal space in the middle layers.
Details
Comparing the middle-layer representations of 8 languages (EN, ZH, AR, RU, JA, KO, HI, FR) across several models, sentences on the same topic clustered closer together regardless of language.
- Sentences about photo synthesis were closer, within Hindi, to Japanese and Chinese sentences than to cooking sentences.
- The same pattern repeated across Qwen3.5-27B, MiniMax M2.5, GLM-4.7, GPT-OSS-120B, Gemma-4 31B.
- In both dense transformer and MoE architectures, language identity weakened toward the middle layers.
In a harder test, the same concept was rendered as an English explanation, a Python function, and a LaTeX formula. As a result, ½mv², 0.5 * m * v ** 2, and "half the mass times velocity squared" showed a tendency to converge on the same region in the internal space.
The conclusion is that these models operate by compressing not just language but the representational format itself into a unified geometric space. The author interprets this in connection with Sapir-Whorf and Chomsky's perspectives, and previews further verification of whether the same phenomenon holds in larger models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.