Tencent open-sources AngelSpec to fix why one-size-fits-all draft models fail
MarkTechPost · Jul 30, 2026 · 2 min read
Speculative decoding has a dirty secret: a single drafter model, no matter how well-tuned on benchmarks, breaks down un...
6 articles
Explore our coverage of speculative-decoding — 6 curated articles, summaries, and related resources from the Inblix archive.
MarkTechPost · Jul 30, 2026 · 2 min read
Speculative decoding has a dirty secret: a single drafter model, no matter how well-tuned on benchmarks, breaks down un...
Hugging Face Blog · Sep 29, 2025 · 2 min read
Intel’s latest experiments are putting the squeeze on draft models to make local AI agents feel less like a waiting gam...
Hugging Face Blog · Nov 20, 2024 · 2 min read
Speculative decoding has been the go-to trick for speeding up large language models, but it always came with an annoyin...
Hugging Face Blog · Oct 29, 2024 · 2 min read
The dirty secret of assisted generation—that speed-boosting technique also called speculative decoding—has always been...
Hugging Face Blog · Oct 8, 2024 · 2 min read
The whole premise of speculative decoding is clever but wasteful. You let a tiny, fast draft model spit out a bunch of...
Hugging Face Blog · Jul 31, 2024 · 2 min read
Google has expanded its Gemma 2 family with a 2.6 billion parameter model designed specifically for on-device deploymen...