clean_stale_shared_memory duplicating the master process when called in a train.py script
Open
Nobody has claimed this yet.
bug
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 206
- PR merge metrics
- No merged PRs in 30d
Description
To reproduce
calling clean_stale_shared_memory() at the beginning of a train.py script itself launched with composer in a distributed setup.
Expected behavior
The memory is cleaned at the beginning of the training and then the training happens normally
What I get:
The process is duplicated on the GPU:0 and is never destroyed
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by tracing clean_stale_shared_memory() from the train.py entry point when launched with Composer in a distributed setup. Reproduce the reported GPU:0 process duplication, then determine why the extra process is not destroyed and verify that shared memory is cleaned before training proceeds normally.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100