String-matching metrics are failing VQA models—even when answers are right
Hugging Face Blog · Jul 25, 2024 · 2 min read
The way we score visual question answering models is quietly falling apart. On Docmatix, a synthetic document VQA datas...
Cutting-edge AI research papers, breakthroughs, and academic developments.
601 articles
Hugging Face Blog · Jul 25, 2024 · 2 min read
The way we score visual question answering models is quietly falling apart. On Docmatix, a synthetic document VQA datas...
Hugging Face Blog · Jul 23, 2024 · 2 min read
Meta just dropped Llama 3.1, and the headline number is absurd: 405 billion parameters. That's the largest dense open-w...
Hugging Face Blog · Jul 22, 2024 · 2 min read
Apple's WWDC 24 Core ML updates make running a 7-billion-parameter model on consumer hardware surprisingly practical. T...
Hugging Face Blog · Jul 18, 2024 · 3 min read
The team behind Idefics2 just released Docmatix, a document visual question answering dataset that makes the previous s...
Hugging Face Blog · Jul 18, 2024 · 2 min read
Hugging Face just solved one of the most annoying problems in LLM deployment: paying for multiple GPUs when you need mu...
Hugging Face Blog · Jul 16, 2024 · 2 min read
The team behind Argilla has open-sourced a practical blueprint for building documentation chatbots that actually unders...
Hugging Face Blog · Jul 16, 2024 · 2 min read
Hugging Face just dropped SmolLM, a family of three small language models ranging from 135M to 1.7B parameters, and the...
Hugging Face Blog · Jul 11, 2024 · 2 min read
The Numina team has just pulled off something remarkable: a fine-tuned 7-billion-parameter model that out-reasoned much...
Hugging Face Blog · Jul 10, 2024 · 2 min read
Hugging Face is rolling out an experimental feature that automatically scans datasets on the Hub for personally identif...
Hugging Face Blog · Jul 10, 2024 · 2 min read
Hugging Face's TRL library just got a capability upgrade that quietly matters more than most model releases: direct pre...
Hugging Face Blog · Jul 10, 2024 · 2 min read
The wall between Hugging Face's massive model collection and KerasHub just came down. Previously, KerasHub users could...
Hugging Face Blog · Jul 9, 2024 · 2 min read
France's Banque des Territoires is betting that generative AI can accelerate one of the country's most ambitious enviro...
Hugging Face Blog · Jul 9, 2024 · 2 min read
Hugging Face developers can now rent Google's custom TPU v5e chips directly through Inference Endpoints and Spaces, wit...
Hugging Face Blog · Jul 8, 2024 · 2 min read
Hugging Face just gave its Dataset Hub a serious upgrade, rolling out four new search filters that make finding the rig...
Hugging Face Blog · Jul 3, 2024 · 2 min read
Intel and MILA have re-architected ProtST, the multi-modal protein language model that made waves at ICML 2023, and rel...
Hugging Face Blog · Jul 1, 2024 · 2 min read
Hugging Face engineers decided to stress-test their Transformers Agents library against GAIA, widely considered the mos...
Hugging Face Blog · Jun 27, 2024 · 2 min read
Google just released Gemma 2, and the numbers are frankly uncomfortable for anyone who bought into the bigger-is-better...
Hugging Face Blog · Jun 25, 2024 · 2 min read
Patent analysis has always been a slog for AI. Generic models choke on claim language, legal jargon, and the kind of te...
Hugging Face Blog · Jun 24, 2024 · 2 min read
Everyone talks about model architecture, parameter counts, and compute budgets. Fewer people want to discuss the unglam...
Hugging Face Blog · Jun 24, 2024 · 2 min read
Florence-2 was supposed to handle visual question answering out of the box. That's what Microsoft's paper implied. But...
Hugging Face Blog · Jun 20, 2024 · 2 min read
Hugging Face's Data Is Better Together initiative started with a simple bet: that open-source AI needs better datasets,...
Hugging Face Blog · Jun 19, 2024 · 3 min read
Prezi, the online presentation platform, has been quietly reworking the machine learning that powers its flagship AI pr...
Hugging Face Blog · Jun 18, 2024 · 2 min read
The AI community has been stuck with benchmarks that either skew too academic or lean too niche. HumanEval is convenien...
Hugging Face Blog · Jun 13, 2024 · 2 min read
Running the same training pipeline with DeepSpeed and PyTorch FSDP should produce similar results. But when Hugging Fac...