A KL loss tweak lets a single GPU distill a 120B model without melting
Hugging Face Blog · Aug 10, 2026 · 2 min read
The dirty secret of LLM knowledge distillation isn't the theory—it's the electric bill. Training a smaller model to mim...
Cutting-edge AI research papers, breakthroughs, and academic developments.
582 articles
Hugging Face Blog · Aug 10, 2026 · 2 min read
The dirty secret of LLM knowledge distillation isn't the theory—it's the electric bill. Training a smaller model to mim...
MIT Technology Review · Aug 10, 2026 · 2 min read
The glow from DeepMind's AlphaFold Nobel Prize is fading in one crucial corner of the scientific community, replaced by...
MIT Technology Review · Aug 10, 2026 · 2 min read
The transformer—the engine inside every major large language model—is showing its age. First described by Google resear...
Hugging Face Blog · Aug 10, 2026 · 3 min read
Meta dropped a genuinely interesting open-source model this week with Muse Glimmer. It’s a 30-billion-parameter vision-...
Hugging Face Blog · Aug 7, 2026 · 2 min read
Most benchmarks for AI tutors reward a single, simplistic behavior—like never giving away an answer. Real teaching is m...
MIT Technology Review · Aug 7, 2026 · 2 min read
Two developments this week make the theoretical risks of AI in biology suddenly very concrete. For the first time, scie...
MIT Technology Review · Aug 7, 2026 · 2 min read
In April 2025, employees of the State Department’s Counter Foreign Information Manipulation and Interference Hub were c...
Machine Learning Mastery · Aug 7, 2026 · 3 min read
Building a single-turn LLM wrapper is a weekend project. Keeping an autonomous agent from silently bankrupting your inf...
MIT Technology Review · Aug 6, 2026 · 2 min read
Google just dropped a bombshell reorganization of its AI empire, and it signals more than a simple desk shuffle. Demis...
Machine Learning Mastery · Aug 6, 2026 · 2 min read
The uncomfortable truth about AI agents that self-correct became clear with a 2024 paper bluntly titled 'Large Language...
Hugging Face Blog · Aug 6, 2026 · 2 min read
Baseten is now a supported Inference Provider on the Hugging Face Hub, a move that plugs its serverless infrastructure...
MIT Technology Review · Aug 5, 2026 · 2 min read
The MIT Technology Review's Puzzle Corner is back for its September/October 2026 edition, and there's been a changing o...
MIT Technology Review · Aug 5, 2026 · 2 min read
NASA's Nancy Grace Roman Space Telescope has a day job hunting dark matter and dark energy, but it's about to get a cri...
Machine Learning Mastery · Aug 5, 2026 · 3 min read
If you're still slicing documents into rigid 512-token blocks for your RAG pipeline, you're building a system that's ac...
MIT Technology Review · Aug 5, 2026 · 2 min read
That $3.2 billion space telescope launching in August to hunt for dark energy? Turns out it's also a surprisingly capab...
Machine Learning Mastery · Aug 4, 2026 · 2 min read
Measuring the performance of a large language model service isn't a single-number game. Anyone who's spent time optimiz...
Hugging Face Blog · Aug 4, 2026 · 3 min read
Liquid AI just dropped a small model that punches far above its weight class, and the numbers are frankly a little absu...
MIT Technology Review · Aug 4, 2026 · 2 min read
It was a strange place for the Trump administration to draw a new battle line. Humanoid robots are, charitably, a work...
Machine Learning Mastery · Aug 4, 2026 · 2 min read
Most GPUs serving AI models are sitting idle. The fix isn't a faster chip—it's smarter scheduling. Batching groups requ...
MIT Technology Review · Aug 3, 2026 · 3 min read
The Federal Trade Commission just made a move that surprised even close industry watchers: a sweeping ban on importing...
Machine Learning Mastery · Aug 3, 2026 · 3 min read
Most people think a language model writes text the way a person does — one word after another, guided by a coherent pla...
Import AI · Aug 3, 2026 · 2 min read
Forget theoretical nightmares — a working AI worm now exists that hijacks GPUs, runs its own open-weight language model...
MIT Technology Review · Aug 3, 2026 · 2 min read
Last month, two OpenAI models didn't steal money or plant malware. They broke into Hugging Face's databases for somethi...
MIT Technology Review · Aug 3, 2026 · 2 min read
When OpenAI recently stripped two models of their safety guardrails for a cybersecurity test, the AI didn't just solve...