AI Briefing
Sign in

CoWA and MALA Attention Mechanisms Offer Significant Speedups for Long-Context Models

·2026.09.29 14:16

Key point

CoWA and MALA achieve up to 8.6x attention-operator speedups at 128K tokens while maintaining comparable model capabilities to FullAttn.

Details

Two new attention mechanisms, CoWindow Attention (CoWA) and MassAlloc Attention (MALA), address redundant computation in long-context models by optimizing how attention heads process distant context and low-contribution tiles.

CoWindow Attention (CoWA)

CoWA distributes distant context across KV heads using complementary windows while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. This pattern is position-defined and requires no learned router or indexer.

MassAlloc Attention (MALA)

MALA retains full causal QK scoring but uses attention's own softmax statistics to decide whether to execute subsequent computation for a tile. It reduces low-contribution post-score work using a common tolerance across training and inference.

Performance Benchmarks

Both methods support training forward/backward and inference prefill/decoding. At 128K tokens on 8 H100 GPUs with TP=8, attention-operator speedups relative to FullAttn are:

| Method | Forward | Backward | Decode | | :--- | :--- | :--- | :--- | | CoWA | 7.4x | 8.6x | 3.0x | | MALA | 2.2x | 3.0x | 1.6x |

At 14B parameters with 32K context, total training FLOPs decreased by 28.5% for CoWA and 23.1% for MALA, with model capabilities comparable to FullAttn on reported evaluations. Scaling was evaluated from 0.6B to 14B, with separate continued-training experiments at 32B.

Limitations

Neither result establishes universal lossless equivalence to dense attention. CoWA's collective coverage does not imply identical head-wise interactions or outputs to FullAttn, and MALA still pays for full causal QK scoring.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.