About bugs in SyncBn
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 9k
- Forks
- 1.5k
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 3
Description
Hello, I think there are bugs in SyncBn when the input size between GPUs is different. The same issue in pytorch1.1 is here issue#22192. And someone fix it here. So I think the SyncBn in apex should also fix the bug. Could someone try it?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the SyncBn implementation in Apex and compare its behavior with PyTorch issue #22192 and commit 29ec4769bbcf545e8727184b413f3b4b6a002b44. Reproduce the failure with different input sizes across GPUs, then verify that SyncBn handles those inputs correctly.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100