facebookresearch / facebookresearch/SlowFast

Multigrid Training Time

Open
#292 9 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
7.4k
Forks
1.3k
PR merge metrics
No merged PRs in 30d

Description

Hi,
I'm trying to reproduce the multigrid training on kinetics 400 in more or less 5 hours. I'm using a cluster with 4 V100 per node, SLOWFAST_8x8_R50_stepwise_multigrid.yaml config file with NUM_GPUS 4, NUM_SHARDS 32 and tried some tuning on batch size, workers. While the accuracy seems fine the training time is far from 5 hours even using 256 GPUs.
Any advice would be really welcome.

Thank you for sharing your code.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.