lm-sys / lm-sys/FastChat

Inference does not stop with Llama-2-13B-GPTQ with exllama

Open
#2,721 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
39.5k
Forks
4.8k
PR merge metrics
No merged PRs in 30d

Description

I have been using the combination of running Llama-2-13B-GPTQ with exllama. However, I noticed that the model inference does not stop and it keeps on generating. I was looking at the ouputs in the gradio demo. Has anyone faced similar issue and knows a solution about this.

image

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the behavior with Llama-2-13B-GPTQ and exllama in the Gradio demo, capturing the prompt and generation settings. Trace where inference determines that generation should stop, then verify that output terminates at the expected stopping condition.

Written by the indexing model from the issue text.

Assessment

Domain
ai
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.