AI Pulse by Inblix

Topic: speculative-decoding

6 articles

Explore our coverage of speculative-decoding — 6 curated articles, summaries, and related resources from the Inblix archive.

Tencent open-sources AngelSpec to fix why one-size-fits-all draft models fail — Inblix summary
AI News

Tencent open-sources AngelSpec to fix why one-size-fits-all draft models fail

MarkTechPost · Jul 30, 2026 · 2 min read

Speculative decoding has a dirty secret: a single drafter model, no matter how well-tuned on benchmarks, breaks down un...

Intel Prunes Qwen3 Draft Model to Hit 1.4x Speedup on Local AI Agents — Inblix summary
Research

Intel Prunes Qwen3 Draft Model to Hit 1.4x Speedup on Local AI Agents

Hugging Face Blog · Sep 29, 2025 · 2 min read

Intel’s latest experiments are putting the squeeze on draft models to make local AI agents feel less like a waiting gam...

Meta's LayerSkip eliminates the draft model, cutting LLM memory use by up to 50% — Inblix summary
Research

Meta's LayerSkip eliminates the draft model, cutting LLM memory use by up to 50%

Hugging Face Blog · Nov 20, 2024 · 2 min read

Speculative decoding has been the go-to trick for speeding up large language models, but it always came with an annoyin...

Intel and Hugging Face crack 1.5x-2x faster LLM decoding across any model family — Inblix summary
Research

Intel and Hugging Face crack 1.5x-2x faster LLM decoding across any model family

Hugging Face Blog · Oct 29, 2024 · 2 min read

The dirty secret of assisted generation—that speed-boosting technique also called speculative decoding—has always been...

Hugging Face squeezes 52% faster LLM inference by guessing when a draft model is wrong — Inblix summary
Research

Hugging Face squeezes 52% faster LLM inference by guessing when a draft model is wrong

Hugging Face Blog · Oct 8, 2024 · 2 min read

The whole premise of speculative decoding is clever but wasteful. You let a tiny, fast draft model spit out a bunch of...

Google drops Gemma 2 2B: a 2.6B on-device model that punches above its weight — Inblix summary
Research

Google drops Gemma 2 2B: a 2.6B on-device model that punches above its weight

Hugging Face Blog · Jul 31, 2024 · 2 min read

Google has expanded its Gemma 2 family with a 2.6 billion parameter model designed specifically for on-device deploymen...