Research
Meta's LayerSkip eliminates the draft model, cutting LLM memory use by up to 50%
Hugging Face Blog · Nov 20, 2024 · 2 min read
Speculative decoding has been the go-to trick for speeding up large language models, but it always came with an annoyin...