huggingface / huggingface/accelerate
[FSDP] support activation offloading with FSDP
Open
enhancement
feature request
- Dominant language
- Python
- Stars
- 9.9k
- Forks
- 1.5k
- Avg merge
- 5d 2h
- Merged PRs (30d)
- 27
Description
Support whole model activation offloading with FSDP - working in conjunction with activation checkpointing - via
https://github.com/pytorch/pytorch/blob/e9ebda29d87ce0916ab08c06ab26fd3766a870e5/torch/distributed/algorithms/_checkpoint/checkpoint_wrapper.py#L171-L191
As `apply_activation_checkpointing` does not wrap the overall root module, wrapping the overall root module with this could offload activation between layer, thus release more GPU memory. The diff should be small and i am happy to work on this.
Contributor guide
Assessment
This issue has not been assessed yet.