AI Briefing
KOSign in

LIFT Architecture Enables Transformers to Pass Hidden State to Next Token via Teacher Supervision

·2026.10.01 09:00

Key point

Experiments with models from 135M to 1B parameters show LIFT outperforms standard Transformers on language modeling and reasoning tasks under token-matched budgets.

Details

Standard Transformers operate as feed-forward networks where deep-layer representations are never fed back to shallower layers, forcing models to recompute intermediate results and discard alternative continuations. The LIFT (Latent Information Feedback Transformer) architecture removes this bottleneck during pretraining by enabling models to propagate state across generation steps.

How LIFT Works

The method turns recurrent-state learning into a prediction problem by pairing each input token with an information-dense state derived from an off-the-shelf pretrained LM. The model, extended with a small number of additional parameters, is trained to predict both the next token and the next state. Because these input states are precomputed, the pretraining process remains fully parallel across positions.

Performance and Efficiency

During inference, the model uses its own predicted states, introducing a minor computational overhead that decreases as model size increases. Experiments with models ranging from 135M to 1B parameters demonstrate that LIFT consistently outperforms standard Transformers and baselines on downstream reasoning tasks and language modeling under token-matched budgets, while matching or exceeding compute-matched Transformers. Additionally, a controlled study on a state-tracking task showed that a tiny LIFT model outperformed a standard Transformer trained on 8x more data.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.