intel / intel/xFasterTransformer

[Feature] Support Medusa decode (Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads)

Open
#489 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
435
Forks
76
PR merge metrics
No merged PRs in 30d

Description

Motivation
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads (https://github.com/FasterDecoding/Medusa) The proposed method can greatly improve the inference speed

Related resources
https://github.com/FasterDecoding/Medusa

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.