[BUG] Resource temporarily unavailable
- Dominant language
- Python
- Stars
- 450
- Forks
- 40
- PR merge metrics
- No merged PRs in 30d
Description
**Describe the bug**
I'm using megatron to train llama2-7b's rlhf alignment with 8 H100s, 60-core CPU, 1200g memory and got this error: Resource temporarily unavailable
**Screenshots**
If applicable, add screenshots to help explain your problem.
error info


gpu use info

Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the two error screenshots and the reported training setup: Llama2-7B RLHF alignment with 8 H100 GPUs, 60 CPU cores, and 1200 GB memory. Reproduce the failure if possible and trace which component emits “Resource temporarily unavailable”; done means identifying a verified cause and documenting or fixing it.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100