KV Caching from Scratch in nanoVLM Delivers a 38% Generation Speedup
Hugging Face Blog · Jun 4, 2025 · 2 min read
Implementing a fundamental optimization from scratch is often the best way to truly understand it. That's exactly what...
Cutting-edge AI research papers, breakthroughs, and academic developments.
595 articles
Hugging Face Blog · Jun 4, 2025 · 2 min read
Implementing a fundamental optimization from scratch is often the best way to truly understand it. That's exactly what...
Hugging Face Blog · Jun 3, 2025 · 2 min read
A software engineer and music producer has demonstrated that you don’t need a GPU cluster to generate studio-ready audi...
Hugging Face Blog · Jun 3, 2025 · 2 min read
H Company just released Holo1, an open-source family of action VLMs that makes web automation dramatically cheaper with...
Hugging Face Blog · Jun 3, 2025 · 2 min read
Robotics research has a scaling problem, and it's not the one you think. While labs with deep pockets train massive vis...
Hugging Face Blog · Jun 3, 2025 · 2 min read
The latest release of Hugging Face's Transformer Reinforcement Learning (TRL) library, v0.18.0, solves a maddening inef...
Hugging Face Blog · May 28, 2025 · 2 min read
The way AI agents take action is getting a quiet but important upgrade. For years, the default has been JSON function c...
Hugging Face Blog · May 25, 2025 · 2 min read
Another day, another obscure shape mismatch tearing through a perfectly good training run. This time, the culprit is a...
Hugging Face Blog · May 23, 2025 · 3 min read
Dell is making a serious play for the air-gapped enterprise, and they just cut the timeline from weeks to under an hour...
Hugging Face Blog · May 23, 2025 · 2 min read
Hugging Face has a compelling argument for AI skeptics who think building agents requires complex frameworks: you can d...
Hugging Face Blog · May 21, 2025 · 2 min read
The Falcon-H1 family dropped today, and it’s a direct challenge to the idea that bigger models are the only path to bet...
Hugging Face Blog · May 21, 2025 · 2 min read
The Technology Innovation Institute (TII) just dropped Falcon-Arabic, and the benchmark numbers are frankly startling....
Hugging Face Blog · May 21, 2025 · 2 min read
Hugging Face just released nanoVLM, a deliberately minimal toolkit for training vision-language models that feels like...
Hugging Face Blog · May 21, 2025 · 2 min read
Hugging Face just made it much cheaper to run the massive Flux image model locally. A new pipeline-level quantization c...
Hugging Face Blog · May 19, 2025 · 2 min read
Two years ago, Microsoft and Hugging Face began a modest experiment: make open models easier to find and deploy on Azur...
Hugging Face Blog · May 15, 2025 · 2 min read
The Falcon-Edge series tackles the central headache of extreme model compression head-on: you usually have to pick betw...
Hugging Face Blog · May 15, 2025 · 2 min read
The Transformers library isn't just growing. It's quietly becoming the central nervous system for the entire machine le...
Hugging Face Blog · May 14, 2025 · 2 min read
Kaggle just made the two biggest platforms in open-source AI a whole lot closer. Starting today, a new integration lets...
Hugging Face Blog · May 13, 2025 · 2 min read
Hugging Face just dropped a community-powered ASR endpoint that's hard to ignore. The new Whisper deployment, built on...
Hugging Face Blog · May 12, 2025 · 2 min read
The vision language model space has gotten weird — in the best way. When we last checked in on VLMs in April 2024, LLaV...
Hugging Face Blog · May 11, 2025 · 2 min read
The biggest hurdle for generalist robots isn't shaky hands or slow processors. It's data. Specifically, the fact that a...
Hugging Face Blog · Apr 30, 2025 · 2 min read
Gradio, the Python library used by over a million developers monthly to build ML interfaces, just made a quiet but sign...
Hugging Face Blog · Apr 30, 2025 · 3 min read
The new Qwen-3 model from Alibaba's Qwen team ships with a chat template that's a quiet masterclass in practical AI des...
Hugging Face Blog · Apr 29, 2025 · 2 min read
Here's a number that makes other quantization researchers wince: 2.1x higher relative accuracy at 2-bit precision. That...
Hugging Face Blog · Apr 29, 2025 · 2 min read
Meta just released Llama Guard 4, and the headline feature isn't just about better detection — it's about practicality....