abetlen / abetlen/llama-cpp-python

Can RotorQuant/TurboQuant and dflash support be added?

Open
#2,184 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
10.6k
Forks
1.4k
PR merge metrics
PR metrics pending

Description

Hi!
This project is very useful for using with llama.cpp in python. However, it would be great if we could have some features even before the main llama.cpp has full support for them. By that I mean TurboQuant/RotorQuant and dFlash speculative decoding. It would make this entire library AMAZING to use.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.