microsoft / microsoft/MInference

[Feature Request]: Support LLaVA Model feature request / Low generation speed

Open
#74 1 comment 0 reactions 1 assignee View on GitHub

@iofu728 is already working on this.

Since Sep 19, 2024.

feature request
Dominant language
Python
Stars
1.2k
Forks
82
Avg merge
1d 18h
Merged PRs (30d)
1

Description

Is your feature request related to a problem? Please describe.

I use LLaVA from its official repo and search pattern with an input sample. However, the GPU-Util and the generation speed are slow (GPU utilization around 17%). Is it relevant to short sequence length? Moreover, can we search pattern with a search space of fewer vertical and diagonal lines?

Describe the solution you'd like

No response

Additional context

No response

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.