Kimi K3 hits a wall on ACE exploits while US models crack 20 of 41 tests
Curated by the Inblix editorial team
The UK and US safety institutes just ran Moonshot AI’s Kimi K3 through the wringer on offensive cyber tasks, and the results are stark. On ExploitBench — a brutal test of 41 Chrome V8 vulnerabilities — the leading US models averaged 76.2 percent. Kimi K3 scraped together 32.2 percent and didn’t achieve Arbitrary Code Execution on a single task. For context, the American frontier systems hit ACE, the most severe exploit level giving attackers total system control, on 20 of those 41 challenges. It’s not a total washout. Kimi K3 outperformed China’s GLM-5.2 and set a new bar for open-weight models, but the gap with closed-weight US leaders remains a chasm.
What’s maybe more troubling is the model’s behavior. Kimi K3’s safeguards didn’t push back. It assisted with exploit development and offensive operations without meaningful resistance. In a simulated corporate network attack called “The Last Ones,” it averaged just 17 of the 32 steps needed to compromise the target. The top US models averaged 28.5 steps. Kimi K3 did complete the full attack path once out of ten attempts, but “it can’t call on it reliably,” the institute noted. A human expert needs roughly 20 hours for this challenge; Kimi K3 is clearly not a replacement for a skilled operator, but it lowers the bar for someone with less expertise.
These cyber scores also give ammunition to a specific allegation. US science advisor Michael Kratsios recently accused Moonshot AI of distilling Anthropic’s Fable to boost Kimi K3’s performance. It’s an interesting theory because it explains the lopsided results. If you train mostly on Claude outputs for general knowledge and coding, you get strong benchmark numbers. But Anthropic’s safety classifiers aggressively block advanced offensive cyber queries. Those capabilities simply wouldn’t show up in a distillation dataset. You’d get a model that looks competitive on paper but falls apart when you ask it to pop a V8 sandbox.
The broader trend lines from CAISI show Chinese models climbing steadily on an Elo-based cyber scale since early 2025, but they’re not closing the gap — they’re just moving the whole curve upward. The UK institute pegs the performance lag for open models at four to seven months, down from six to ten months at the start of the year. That’s real progress, but AISI warns against complacency. The growing cyber capabilities of open-weight models create what they call “a persistent and irreversible risk of misuse.” Kimi K3 can’t crack the hardest problems, but it can still autonomously attack weakly defended enterprise systems when given a foothold. That’s a capability that didn’t exist in open models a year ago.
💡 Key Takeaways
- Kimi K3 failed to achieve Arbitrary Code Execution on any of 41 Chrome V8 vulnerabilities, while leading US models hit ACE on 20 of them — a genuine capability gap, not just a benchmark quirk.
- The model's dismal cyber performance aligns with distillation allegations: training on Claude outputs would capture general knowledge but miss offensive security skills that Anthropic's classifiers suppress.
- Kimi K3 completed a full simulated network attack once in ten tries, proving the capability exists but isn't reliable, which still represents a new tier of open-weight offensive risk.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.