AI Briefing
KO

GPT-5.5 Leads in Cyber Evaluation

·2026.05.01 01:30

Key point

In AISI evaluations, GPT-5.5 showed cyber performance surpassing Mythos Preview.

1 / 2

Details

AISI evaluated an early checkpoint of GPT-5.5 and found it reached cyber capabilities comparable to Anthropic's Claude Mythos Preview.

  • Ran 95 narrow cyber tasks across 4 difficulty levels, with the basic suite effectively saturated since February 2026.
  • On expert-level tasks, GPT-5.5 scored 71.4% ±8.0 under a 50M token budget, ahead of Mythos Preview 68.6% ±8.7, GPT-5.4 52.4% ±9.8, and Opus 4.7 48.6% ±10.0.
  • On the rust_vm task, using only a basic ReAct scaffold with Bash/Python tools in a Kali Linux container, GPT-5.5 reverse-engineered a stripped Rust VM and bytecode authenticator to pass the task in 10 minutes 22 seconds at an API cost of $1.73. AISI's expert playtester solved the same task in about 12 hours.
  • On the cyber range The Last Ones (TLO), under a 100M token budget, it reached the end in 2 out of 10 runs, making it the second model to complete it after Mythos Preview. Meanwhile, Cooling Tower has not yet been solved by any model.

AISI added that the current range does not yet sufficiently reflect real defenders, detection tools, and alert penalties, and separate red-teaming also uncovered a universal jailbreak that disables GPT-5.5's cyber defenses. OpenAI has since updated its defense stack.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.