deepspeedai / deepspeedai/DeepSpeed
BERT-Large can't fit in VRAM 16G,RAM 64G
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
sequence length 400, batch size per-device 8, fp16,num_gpus 8, one node
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the reported BERT-Large setup: sequence length 400, per-device batch size 8, fp16, eight GPUs on one node, with 16G VRAM and 64G RAM. The issue names no files, tests, logs, versions, or entry point, so first determine where the memory failure occurs and what evidence defines a successful fix.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100