Repeated tokens generated from 'generate.py' running on GPU
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 40.3k
- Forks
- 3.7k
- PR merge metrics
- No merged PRs in 30d
Description
Dear Authors,
Thanks for introducing the amazing project. When I tested the BitNet Inference Kernel on RTX 3090 with Ubuntu system, I followed the commands in README.md, but I got repeated tokens as the output. For example:
Could you help me explain Python?
OfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOfOf
Could you help me check if anything could be wrong here? Thanks.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the commands in README.md and the generate.py entry point, reproducing the repeated-token output on an RTX 3090 with Ubuntu. Compare the generated output with the expected inference behavior and document or correct the cause so the prompt no longer produces repeated tokens.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp, python, ubuntu
- Domain
- backend, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100