LLM Encrypted Reasoning Leaked
Key point
Researchers reported an attack that decrypts and leaks encrypted reasoning records from proprietary LLMs.
Details
Researchers reported that encrypted chain-of-thought blocks returned to clients by the APIs of Anthropic, OpenAI, and Google can be reused across different sessions, users, and models.
By exploiting the fact that members of the same model family use the same encryption key, they injected reasoning records generated by a stronger model into a weaker model and jailbroke it to output the original reasoning. Claude Haiku 4.5 was identified as the most vulnerable model.
The paper also presented a prompt injection variant where malicious instructions such as data exfiltration are embedded in a model's reasoning record and then reinjected into another model. The three companies stated that they fixed the issue to prevent the attack from being reproduced after receiving the researchers' report.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.