Research
Intel and Hugging Face crack 1.5x-2x faster LLM decoding across any model family
Hugging Face Blog · Oct 29, 2024 · 2 min read
The dirty secret of assisted generation—that speed-boosting technique also called speculative decoding—has always been...