AI4Finance-Foundation / AI4Finance-Foundation/RLSolver
✨ DataParallel and DistributedDataParallel for speed up training.
- Dominant language
- Python
- Stars
- 169
- Forks
- 36
- PR merge metrics
- No merged PRs in 30d
Description
DataParallel: multiple thread for single machine multiple GPUs
- unbalance GPU memory and GPU usage. ([discuss.pytorch.org: Use `FullModel` which writes loss function into the model to solve the memory usage imbalance problem. ](https://discuss.pytorch.org/t/dataparallel-imbalanced-memory-usage/22551/6))
- slow
- Collecting gradients by a serial method
DistributedDataParallel: multiple processing for single or multiple machines and multiple GPUs.
- balance GPU memory and GPU usage. (don't need to use `FullModel`)
- faster than DataParallel
- [Ring-Allreduce by pytorch](https://pytorch.org/tutorials/intermediate/dist_tuto.html#our-own-ring-allreduce)
It is very easy to add **DataParallel** into the code, but DataParallel brings less speed up.
It's a little tricky to use because **DistributedDataParallel** needs to be started from the command line, but it gives a significant speedup with 4 GPUs in single machine in high GPU memory.
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, entry points, or tests are identified in the issue. Start by locating the training code and determining how GPU workers are currently launched, then compare the required DataParallel or DistributedDataParallel integration with the existing training flow. Done should include the selected multi-GPU approach, documented launch usage, and evidence of the intended speed or memory improvement.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100