Scanning 500K Hugging Face models found a 'dangerous middle' nobody talks about
Curated by the Inblix editorial team
There are over half a million models on the Hugging Face hub, and until now, picking one felt like a gamble on safety. A new initiative called RiskRubric.ai, spearheaded by the Cloud Security Alliance and Noma Security with help from Haize Labs and Harmonic Security, just dropped a standardized grading system to cut through the noise. They’re hitting models with 1,000+ reliability tests and 200+ adversarial security probes, then slapping a clean A-F letter grade on them.
The initial scan, which looked at both open and closed models, revealed a polarized landscape. The scores range from a dismal 47 to a stellar 94 out of 100, with a median of 81. While 54% of models scored an A or B, a long tail of underperformers is dragging the average down. Noma Security’s data shows a cluster of models stuck in the 50–67 band—the C and D range. These aren’t outright broken, but they’re not safe either. “This band represents the most practical area of concern, where security gaps are material enough to warrant prioritization,” the report notes. It’s the perfect hunting ground for attackers.
The researchers also found a tight coupling between security and safety. Models that hardened their defenses against prompt injections and jailbreaks almost always scored better on preventing harmful outputs. Security isn’t just about blocking hacks; it directly reduces downstream harms. But there’s a catch. Stricter guardrails often kill transparency, making models refuse requests without any explanation. That creates an opaque, trust-eroding experience even when the system is technically secure.
Interestingly, open models are shining where you’d expect: transparency. They’re consistently outscoring closed-source counterparts because you can actually see the training data and methods. The platform lets you filter by what you care about—privacy scores for healthcare, reliability for customer-facing apps—so you don’t have to trade off one risk for another. The goal is a virtuous cycle where public, standardized audits force everyone to patch their weakest links. The real question now is whether enterprise procurement teams will actually start requiring a B+ before deployment, or if the dangerous middle keeps getting a free pass.
💡 Key Takeaways
- A full 54% of evaluated models scored an A or B in risk assessment, but a concerning cluster in the 50–67 range represents a material security gap attackers are likely to exploit.
- Investing in core security controls like prompt injection defenses directly reduces harmful outputs, proving safety is largely a byproduct of robust security posture.
- Stricter safety guardrails frequently erode transparency by generating unexplained refusals, forcing a trade-off between user trust and raw security hardening.
- Open models consistently outperform closed models on transparency metrics because their development practices allow for comprehensive documentation review.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.