Research
Hugging Face packs 2x training speed bump into Flash Attention 2 with a single collator swap
Hugging Face Blog · Aug 21, 2024 · 2 min read
Padding tokens have always been the silent performance killer in LLM training. You batch your examples, stuff them with...