AI can't do your data analyst job yet: Hugging Face benchmark scores just 16%
Curated by the Inblix editorial team
If you’re a data analyst waiting for AI to take the boring parts of your job off your plate, you’re going to be waiting a while longer. A brutally honest new benchmark from Hugging Face and the payments giant Adyen shows that even the most advanced AI agents crash and burn on real-world analytical work. The best reasoning systems they tested hit a ceiling of just 16% accuracy on DABstep, a set of over 450 genuine tasks pulled directly from Adyen’s actual workloads.
These aren’t textbook SQL puzzles or sanitized Kaggle competitions. DABstep forces models to grapple with the messy reality that analysts face daily: balancing structured databases against free-form text and scattered documentation, connecting queries to real business use-cases, and doing it all without hallucinating numbers that could steer a company off a cliff. As the researchers point out, data analysis is rarely a clean, linear process. It requires juggling technical plumbing, domain-specific context, and iterative reasoning — exactly the kind of cognitive load that breaks brittle AI systems.
What makes this benchmark sting is its simplicity. It doesn’t require complex configurations like SWE-bench. Models just get a code execution environment and the data. Evaluation is strictly factoid and binary — right or wrong, no room for a smooth-talking LLM to argue its way to partial credit. This exposes a deep gap between the impressive demos of agentic workflows and the precision required when the output isn’t a poem but a business recommendation. The challenge is in navigating ambiguity across multiple datasets and documents without the safety net of single-question isolation.
The 16% score is a cold splash of water on the idea that we’re close to deploying autonomous data analysts. While agentic systems have shown teeth in software engineering and open-ended QA, DABstep proves that the combination of rigorous multi-step reasoning and intolerance for hallucination remains a massive unsolved problem. For now, the messy, creative, and tedious work of analysis is still very much a human game. It’s a stark reminder that benchmarks testing isolated code generation aren’t measuring the same thing as actual job performance.
💡 Key Takeaways
- The top AI agent scored only 16% on DABstep, a benchmark built from real Adyen analyst tasks, proving current models fail at multi-step, hallucination-free data work.
- DABstep's binary evaluation leaves no room for AI smooth-talking — answers are objectively right or wrong, which exposes the fragility of agentic reasoning on complex, unstructured data.
- Unlike Kaggle or isolated SQL tests, these tasks force models to connect free-form text, structured databases, and actual business context, mimicking the cognitive load of a real analyst.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.