intel / intel/xFasterTransformer
[Feature] Support Medusa decode (Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads)
Open
- Dominant language
- C++
- Stars
- 435
- Forks
- 76
- PR merge metrics
- No merged PRs in 30d
Description
Motivation
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (https://github.com/FasterDecoding/Medusa) The proposed method can greatly improve the inference speed
Related resources
https://github.com/FasterDecoding/Medusa
Contributor guide
Assessment
This issue has not been assessed yet.