microsoft / microsoft/MInference
[Feature Request]: Support LLaVA Model feature request / Low generation speed
Open
@iofu728 is already working on this.
Since Sep 19, 2024.
feature request
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 82
- Avg merge
- 1d 18h
- Merged PRs (30d)
- 1
Description
Is your feature request related to a problem? Please describe.
I use LLaVA from its official repo and search pattern with an input sample. However, the GPU-Util and the generation speed are slow (GPU utilization around 17%). Is it relevant to short sequence length? Moreover, can we search pattern with a search space of fewer vertical and diagonal lines?
Describe the solution you'd like
No response
Additional context
No response
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.