Inference does not stop with Llama-2-13B-GPTQ with exllama
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
I have been using the combination of running Llama-2-13B-GPTQ with exllama. However, I noticed that the model inference does not stop and it keeps on generating. I was looking at the ouputs in the gradio demo. Has anyone faced similar issue and knows a solution about this.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the behavior with Llama-2-13B-GPTQ and exllama in the Gradio demo, capturing the prompt and generation settings. Trace where inference determines that generation should stop, then verify that output terminates at the expected stopping condition.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100