deepspeedai / deepspeedai/DeepSpeed
AutoModelFromCausalLLM of Bloom not releasing GPU memory after each inference batch [BUG]
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Describe the bug
Hi there, I have set torch.no_grad() and torch.cuda.empty_cache(), but the GPU still encounters out-of-memory (OOM) errors after a few inferences. My torch version is 1.13.1, deepspeed version is 0.9, and transformer version is 4.28, cuda driver 11.6 with v100
Expected behavior
auto release the memory
ds_report output
all good
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No file or test is named; start at the AutoModelFromCausalLLM Bloom inference entry point and reproduce the reported GPU memory growth across batches with the stated PyTorch, DeepSpeed, Transformers, and CUDA versions. Compare allocated memory after each batch and define done as repeated inference without progressive growth or an OOM.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100