abetlen / abetlen/llama-cpp-python
Can RotorQuant/TurboQuant and dflash support be added?
未關閉
- 主要語言
- Python
- 星號
- 10.6k
- 分支
- 1.4k
- PR 合併指標
- PR 指標待擷取
描述
Hi!
This project is very useful for using with llama.cpp in python. However, it would be great if we could have some features even before the main llama.cpp has full support for them. By that I mean TurboQuant/RotorQuant and dFlash speculative decoding. It would make this entire library AMAZING to use.
貢獻指南
評估
這個 Issue 還沒有評估資料。