lllyasviel / lllyasviel/ControlNet

RuntimeError when training on multiple GPUs

Open
#202 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
34.1k
Forks
3k
PR merge metrics
No merged PRs in 30d

Description

I tried to train on multiple GPUs, but when reading the data, even if I set num_workers=0, I still get the error
RuntimeError: unable to open shared memory object
and I don't have root access, so I can't increase the openfile data.
trainer = pl.Trainer(gpus=2, precision=32, callbacks=[logger])
As soon as I change gpus to 1, training works fine. Anyone have ideas?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the reported Trainer configuration with gpus=2, precision=32, num_workers=0, and the logger callback, then compare it with the working single-GPU case. Trace the multi-GPU data-reading path and shared-memory failure; done means multi-GPU training works without requiring root access or increased open-file limits.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.