AI Pulse by Inblix

EVMbench tests AI agents on smart contract security

OpenAI Blog · Jul 10, 2026 · 1 min read · Read original article →

Curated by the Inblix editorial team


Smart contracts hold over $100 billion in crypto assets, making their security paramount. Now, a new benchmark called EVMbench is putting AI agents to the test. Created with Paradigm, it evaluates whether AI can detect, patch, or exploit real-world vulnerabilities in Ethereum-style blockchain environments. The benchmark covers 117 high-severity bugs from 40 audits, including scenarios from the Tempo blockchain for stablecoin payments. AI agents are scored in three modes: finding vulnerabilities (detect), fixing them without breaking functionality (patch), and actually draining funds in a sandbox (exploit). The best model, GPT‑5.3‑Codex, scored 71% in exploit mode—a big leap over earlier attempts. To prevent cheating, exploits are run in an isolated environment and graded programmatically. Why it matters: As AI gets better at coding, it can either become a hacker’s weapon or a security guard—EVMbench helps steer it toward the latter.

💡 Key Takeaways

  1. EVMbench is a new benchmark that tests AI agents on detecting, patching, and exploiting smart contract vulnerabilities.
  2. It uses 117 curated vulnerabilities from real audits, making it highly practical for blockchain security.
  3. The best AI model scored 71% in exploit mode, showing rapid improvement in automated attack capabilities.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

← Back to all articles