EVMbench tests AI agents on smart contract security
Curated by the Inblix editorial team
Smart contracts hold over $100 billion in crypto assets, making their security paramount. Now, a new benchmark called EVMbench is putting AI agents to the test. Created with Paradigm, it evaluates whether AI can detect, patch, or exploit real-world vulnerabilities in Ethereum-style blockchain environments. The benchmark covers 117 high-severity bugs from 40 audits, including scenarios from the Tempo blockchain for stablecoin payments. AI agents are scored in three modes: finding vulnerabilities (detect), fixing them without breaking functionality (patch), and actually draining funds in a sandbox (exploit). The best model, GPT‑5.3‑Codex, scored 71% in exploit mode—a big leap over earlier attempts. To prevent cheating, exploits are run in an isolated environment and graded programmatically. Why it matters: As AI gets better at coding, it can either become a hacker’s weapon or a security guard—EVMbench helps steer it toward the latter.
💡 Key Takeaways
- EVMbench is a new benchmark that tests AI agents on detecting, patching, and exploiting smart contract vulnerabilities.
- It uses 117 curated vulnerabilities from real audits, making it highly practical for blockchain security.
- The best AI model scored 71% in exploit mode, showing rapid improvement in automated attack capabilities.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.