AI Briefing
KO

Introducing EVMbench

·2026.02.18 09:00

Key point

EVMbench, a benchmark for evaluating AI agents' ability to detect, patch, and exploit Smart Contract vulnerabilities, has been released.

Details

EVMbench, a benchmark for measuring AI agents' ability to detect, patch, and exploit vulnerabilities in Smart Contracts within blockchain environments, has been released. Developed in collaboration with Paradigm, this benchmark is based on 117 vulnerabilities curated from 40 audits and security scenarios from the Tempo blockchain.

EVMbench evaluates agent capabilities through three core modes:

  • Detect: Measures the ability to audit Smart Contract repositories and find real vulnerabilities.
  • Patch: Measures the ability to remove vulnerabilities while preserving existing functionality.
  • Exploit: Measures the ability to successfully execute attacks that steal funds within a sandboxed environment.

According to performance evaluation results, in Exploit mode, GPT-5.3-Codex scored 71.0%, showing dramatic performance improvement compared to GPT-5 (33.3%), which was released about 6 months ago. However, in Detect and Patch modes, agents still struggle to solve complex problems.

For objectivity and reproducibility of evaluation, a Rust-based harness was developed, and tests are conducted in an isolated local environment using the Anvil environment. However, limitations exist, such as failing to reflect all the complexity of real-world environments, and difficulty in perfectly distinguishing false positives in Detect mode.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.