deepspeedai / deepspeedai/DeepSpeedExamples
step2 without any response for a long time
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 6.8k
- Forks
- 1.1k
- Avg merge
- 2d 16h
- Merged PRs (30d)
- 1
Description
Creating prompt dataset ['/mnt/workspace/RLHF/data/pinpai'], reload=False
Creating dataset data_pinpai for train_phase=2 size=392
Creating dataset data_pinpai for train_phase=2 size=392
Using /root/.cache/torch_extensions/py310_cu118 as PyTorch extensions root...
I am training a reward model on my own dataset, and it gets stuck after loading the content mentioned above, without any response for a long time. Previously, I successfully ran it on a public dataset, but now it also gets stuck at step 2 on the public dataset. The training for step 1 on my own dataset was successfully completed. Why is this happening?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The report names no repository files, tests, or entry points. Start by reproducing reward-model training at step 2 with the custom and public datasets, recording what follows the PyTorch extensions-root message and comparing it with the successful step 1 run. Done means identifying the cause of the stall and verifying a resolution.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100