CacheBack enables receiver-conditioned latent communication for multi-agent systems
Key point
Using Qwen3-8B on FanOutQA, CacheBack achieves 3.2× faster median task completion and 14.7 percentage points higher accuracy than text-based communication.
Details
CacheBack introduces receiver-conditioned latent communication, allowing multi-agent systems to share selected subsets of internal state rather than full contexts or text messages. This approach addresses the bottleneck where agents reading large, separate contexts lose efficiency when communicating via text or retaining every position in their cache.
Performance and Efficiency
In tests using Qwen3-8B on the FanOutQA benchmark, CacheBack demonstrated significant improvements over same-size text communication:
- 3.2× faster median task completion.
- 14.7 percentage points higher accuracy.
- At 16× compression, the method removes approximately 94% of sender positions while still improving accuracy and latency across all tested model families.
The system operates without training, using attention to the receiver’s request to select relevant parts of the sender's already-computed state. Improvements were observed across four model families, including dense Transformers, Mamba-attention hybrids, and sliding-window attention.
Practical Application
A demo involving seven Qwen3-8B workers fixing a Django bug showed a 4.41× speedup (26 seconds vs. 113 seconds) compared to text-based agents, with both runs producing identical patches that passed all 88 tests. The code is open source and supports matching dense Qwen3 models via Hugging Face and vLLM.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.