FactorEngram: Factorized N-gram Memory with Basis-Level Gating Improves Language Model Performance
Key point
The new FactorEngram architecture improves language modeling on 340M and 1B parameter backbones by using basis-level contextual gating for n-gram memory.
Details
FactorEngram introduces a factorized n-gram memory with basis-level contextual gating to address limitations in existing lookup-based memory systems like Engram. Traditional designs treat retrieved embeddings as monolithic units modulated by a single scalar gate, preventing polysemous patterns from selectively accessing relevant memory components and limiting parameter sharing to unrelated hash collisions.
How It Works
FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, allowing related patterns to reuse common components. The architecture uses the same dictionary for gating, where the backbone hidden state is scored against each basis vector to modulate the corresponding coefficient before reconstruction. This enables the context to control each memory component individually.
Performance and Configuration
The method covers both individual tokens and multi-token n-grams. Testing on 340M and 1B-parameter Transformer backbones showed improvements in language modeling and downstream task performance. Ablation studies identified that inserting the memory branch before the attention sublayer in the middle layers is an effective configuration.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.