Holo3.1 ships quantized agents that run locally on your laptop
Hugging Face Blog · Jun 2, 2026 · 2 min read
Hcompany just dropped Holo3.1, and it's not a minor point release. This is the first time the company has shipped quant...
Cutting-edge AI research papers, breakthroughs, and academic developments.
587 articles
Hugging Face Blog · Jun 2, 2026 · 2 min read
Hcompany just dropped Holo3.1, and it's not a minor point release. This is the first time the company has shipped quant...
Hugging Face Blog · Jun 1, 2026 · 2 min read
JetBrains just open-sourced Mellum2, a 12-billion-parameter Mixture-of-Experts model built for the unglamorous, high-fr...
Hugging Face Blog · Jun 1, 2026 · 2 min read
The narrative around enterprise AI has become a broken record: companies launch pilots, then watch them crash into the...
Hugging Face Blog · May 29, 2026 · 2 min read
Most PyTorch users know they should profile their models. Few actually do it. The barrier isn't a lack of tools — it's...
Hugging Face Blog · May 27, 2026 · 2 min read
Every async RL library has a dirty secret: every single step, the trainer ships the entire model to the inference engin...
Hugging Face Blog · May 27, 2026 · 2 min read
A new tutorial from Hugging Face shows how to run Reachy Mini, the open-source robot, completely offline. No cloud. No...
Hugging Face Blog · May 25, 2026 · 2 min read
If you've ever nodded along while secretly wondering what separates an LLM 'harness' from its 'scaffold,' you're in cro...
Hugging Face Blog · May 19, 2026 · 2 min read
The team behind OlmoEarth just made a move that's more about smart engineering than raw power, and it could quietly res...
Hugging Face Blog · May 19, 2026 · 2 min read
A new family of six open-source cross-encoder rerankers, all built on the ModernBERT architecture, is now available and...
Hugging Face Blog · May 18, 2026 · 2 min read
The team behind PaddleOCR just shipped version 3.5, and the real news isn't a new model—it's a new inference engine opt...
Hugging Face Blog · May 14, 2026 · 3 min read
IBM just flipped the script on multilingual embeddings with Granite Multilingual R2. For years, the trade-off was bruta...
Hugging Face Blog · May 14, 2026 · 2 min read
If you're running inference on an H200 at $5 an hour, every second the GPU sits idle is money you're lighting on fire....
Hugging Face Blog · May 11, 2026 · 2 min read
If you're training or running inference on large models, you already know the pain. The step time that should be domina...
Hugging Face Blog · May 6, 2026 · 2 min read
Nobody should have to debug a reinforcement learning pipeline by staring at a clip-rate chart that looks like a heart a...
Hugging Face Blog · May 6, 2026 · 2 min read
The Open ASR Leaderboard just got its first taste of private test data — and it's a direct shot at the subtle art of 'b...
Hugging Face Blog · Apr 29, 2026 · 2 min read
IBM just dropped Granite 4.1, and the playbook they're sharing is a direct challenge to the “scale is all you need” ort...
Hugging Face Blog · Apr 29, 2026 · 2 min read
Hugging Face just got a serious cost-performance boost. DeepInfra, the serverless inference platform known for aggressi...
Hugging Face Blog · Apr 28, 2026 · 2 min read
NVIDIA is taking a swing at the omni-modal crown with the new Nemotron 3 Nano Omni, and the spec sheet suggests it isn'...
Hugging Face Blog · Apr 27, 2026 · 3 min read
The Privacy Filter model isn't another massive, resource-hungry beast. OpenAI shipped it as a 1.5B-parameter Apache 2.0...
Hugging Face Blog · Apr 24, 2026 · 2 min read
The dirty secret of AI agents is that they break. Not dramatically, but in predictable, boring ways that make long-runn...
Hugging Face Blog · Apr 23, 2026 · 2 min read
Building a Chrome extension that runs local AI models isn't just about picking the right library—it's an architectural...
Hugging Face Blog · Apr 21, 2026 · 2 min read
If you've been tracking Arabic LLM evaluation, you've probably noticed a growing tension: the number of benchmarks and...
Hugging Face Blog · Apr 21, 2026 · 3 min read
The AI model getting all the attention is called Mythos, a frontier LLM that’s been finding and fixing software vulnera...
Hugging Face Blog · Apr 16, 2026 · 3 min read
Code agents are working. Jensen Huang says we've gone from 30 million coders to a billion overnight. Anyone with an age...