Hugging Face's open data drive pulled 385 volunteers in days
Curated by the Inblix editorial team
Hugging Face’s Data Is Better Together initiative started with a simple bet: that open-source AI needs better datasets, and that the community wants to build them. The response proved the bet right. Within days of launching the prompt ranking project, more than 385 people signed up to help create a dataset of 10,000 prompts — both synthetic and human-generated — ranked by quality. That effort produced the DIBT/10k_prompts_ranked dataset, now available for prompt ranking tasks or synthetic data generation. It’s already been used to train new models, including SPIN.
The early momentum exposed a blind spot. English-centric data wasn’t going to cut it, and open LLMs lacked solid language-specific benchmarks. So the team launched the Multilingual Prompt Evaluation Project, selecting 500 high-quality prompts from the original dataset for translation. More than 18 language leaders stepped up to create translation spaces, with Dutch, Russian, and Spanish already completed and others in progress. A Discord community of dataset builders has grown around the effort.
The initiative has also produced practical cookbooks for domain-specific datasets, DPO/ORPO datasets, and KTO datasets — guides aimed at helping engineers and domain experts bootstrap their own training data. The team’s honest assessment: the community is eager, but real gaps remain. Certain languages, domains, and tasks are still underrepresented in open-source datasets, and closing those gaps requires more than enthusiasm. It requires tools and documentation that lower the barrier to entry.
For anyone who’s watched open-source models struggle against proprietary ones, this is the unglamorous work that actually matters. Better datasets, not just bigger models, drive progress. The invitation is open: join the #data-is-better-together channel on Hugging Face Discord and contribute. The hard part isn’t the technology — it’s coordination, and that’s already underway.
💡 Key Takeaways
- Hugging Face's DIBT initiative produced a 10,000-prompt ranked dataset with contributions from over 385 volunteers in just days.
- The Multilingual Prompt Evaluation Project is translating 500 high-quality prompts into more than 18 languages to address the lack of non-English benchmarks for open LLMs.
- The DIBT/10k_prompts_ranked dataset has already been used to train new models, including SPIN, proving community-built data can feed real model development.
- Hugging Face acknowledges that certain languages, domains, and tasks remain underrepresented in open-source datasets and is prioritizing tools and documentation to close those gaps.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.