AI Briefing
KO

Microsoft's Multi-Agent AI System Outperforms Anthropic Mythos

·2026.05.14 09:00

Key point

Microsoft's MDASH surpassed Anthropic Mythos on the CyberGym benchmark.

Details

Microsoft unveiled MDASH, delivering results on the CyberGym benchmark that surpass Anthropic's Mythos. MDASH is built on a structure where over 100 specialized AI agents work across multiple models to find vulnerabilities, challenge each other, and verify findings with PoC attacks.

This week, the system discovered 16 new Windows vulnerabilities, 4 of which were critical remote code execution flaws that were patched in this month's Patch Tuesday. Microsoft stated that MDASH is already being used by its internal security engineering team, and will also enter a limited private preview.

On CyberGym, MDASH scored 88.45% to take first place, followed by Mythos Preview at 83.1% and GPT-5.5 at 81.8%. However, all of these scores are self-reported by the respective companies, and none have been independently verified.

The benchmark consists of 1,507 tasks, testing whether the system can produce a working exploit given a description of a known vulnerability and an unpatched codebase. While the performance gains show promise on the defensive side, they are also raising concerns that AI could be used as an offensive hacking tool.

  • MDASH: a system based on a multi-model agentic scanning harness
  • Mythos: Anthropic's single-model agent framework
  • CyberGym: a benchmark that evaluates the ability to reproduce real-world vulnerabilities

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.