facebookresearch / facebookresearch/SlowFast
Multigrid Training Time
Open
- Dominant language
- Python
- Stars
- 7.4k
- Forks
- 1.3k
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
I'm trying to reproduce the multigrid training on kinetics 400 in more or less 5 hours. I'm using a cluster with 4 V100 per node, SLOWFAST_8x8_R50_stepwise_multigrid.yaml config file with NUM_GPUS 4, NUM_SHARDS 32 and tried some tuning on batch size, workers. While the accuracy seems fine the training time is far from 5 hours even using 256 GPUs.
Any advice would be really welcome.
Thank you for sharing your code.
Contributor guide
Assessment
This issue has not been assessed yet.