AI Briefing
KO

Mythos Cyber Shock

·2026.04.16 14:21

Key point

Claude Mythos Preview surpassed previous frontier models by one tier in cyber capability.

1 / 2

Details

AISI conducted a cyber evaluation of Anthropic's Claude Mythos Preview and found it performed one tier higher than existing frontier models.

In CTF, performance improved from technical-novice/apprentice-level tasks up to expert-level tasks, achieving a 73% success rate on expert-level tasks.

In The Last Ones (TLO), a more complex multi-step cyber-attack simulation, it completed an average of 22 steps out of 32 steps total, and 3 out of 10 attempts completed from start to finish. The comparison model, Claude Opus 4.6, only reached an average of 16 steps.

However, it failed to complete the operational technology (OT)-focused Cooling Tower range, getting stuck in the IT segment. AISI explained that since the environment lacked defensive equipment and active response, no definitive conclusions can be drawn about its ability to attack systems that are actually well-defended.

In conclusion, Mythos Preview has reached a level where it can autonomously attack weakly secured enterprise systems once network access is obtained, and it was assessed that model performance could continue to scale further with more inference compute and larger token budgets. Accordingly, organizations should strengthen basic defenses such as patch management, access control, security configuration, and logging.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.