Benchmark Finds LLMs Read Cablese Records at Parity or Better Than Plain Text
Key point
A new benchmark, the Telegraph Test, demonstrates that LLMs can read compressed 'cablese' records with accuracy equal to or better than plaintext records, offering 25–49% token savings for machine-to-agent communication.
Details
Benchmark Overview and Key Findings
The Telegraph Test evaluates how Large Language Models (LLMs) handle cablese, a terse, telegraph-inspired register that drops articles and filler while preserving facts. The study reveals that when LLMs read compressed records (summaries) rather than generating compressed answers directly, comprehension remains at parity or improves compared to plaintext records.
Key findings from the cross-family matrix include:
- Token Savings: Savings vary by model. qwen3.8-27b achieved 48.9% token savings, gemma-4-31b saved 40.4%, and GLM-5.3-Flash achieved 48.4% savings when using lowercase instructions.
- Recovery Ratios: Downstream models reading these compressed records achieved recovery ratios between 0.99 and 1.10 (where 1.00 equals the plaintext record control accuracy). For example, gemma-4-31b and qwen3.8-27b as readers achieved ratios of 1.09 and 1.10 respectively, indicating they understood the compressed records better than the plaintext record control.
- Baseline Context: The plaintext record control accuracy was approximately 75.8% (relative to the full passage baseline of 91.5%), not 100%, due to strict grading and ambiguous questions. The recovery ratios are relative to this plaintext record baseline.
Operational Implications
The research indicates that cablese is effective for agent-to-agent communication and memory storage where the consumer is a model.
- Cost Efficiency: Output tokens cost 3–5x more than input tokens. Compressing machine-consumed text can significantly reduce API bills. Expanding records back to plaintext for human review reduces the net savings by about a third, but the archive remains cheaper than plain text.
- Model Compatibility: The technique works across families like gemma, qwen, and nemotron because the register is latent in training data from historical telegraph and SMS usage.
- Limitations: Models with mandatory reasoning, such as gpt-5-mini, incur higher costs because the terse format triggers more internal thinking, doubling the write cost despite shorter output. The author advises using this technique only with models where reasoning can be disabled or controlled.
Comparison with Other Methods
The article contrasts cablese with other compression techniques:
- Codebooks: Substitution-based methods yield only ~10% savings.
- Emergent Protocols: Methods like BabelTele or GLOSSOGEN can achieve higher compression (up to 6x) but result in protocols that are unstable and illegible to humans.
- Cablese: Offers a middle ground with 25–49% compression (depending on the model), deterministic behavior, and human auditability.
The benchmark is open-sourced at github.com/Travis42/telegraph-test.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.