Research
BigCodeBench stress-tests LLMs with 1,140 tasks and 99% branch coverage
Hugging Face Blog · Jun 18, 2024 · 2 min read
The AI community has been stuck with benchmarks that either skew too academic or lean too niche. HumanEval is convenien...