abetlen / abetlen/llama-cpp-python
Add batched inference
未關閉
enhancement
high-priority
- 主要語言
- Python
- 星號
- 10.6k
- 分支
- 1.4k
- PR 合併指標
- PR 指標待擷取
描述
- [x] Use `llama_decode` instead of deprecated `llama_eval` in `Llama` class
- [ ] Implement batched inference support for `generate` and `create_completion` methods in `Llama` class
- [ ] Add support for streaming / infinite completion
貢獻指南
評估
這個 Issue 還沒有評估資料。