pytorch / pytorch/vision

Question about torchvision.io.decode_image

Open
#4,325 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

awaiting response module: io needs reproduction
Dominant language
Python
Stars
17.9k
Forks
7.3k
Avg merge
1d 15h
Merged PRs (30d)
13

Description

when we use torchvision.io.decode_image(img,device = local_rank) to train with ddp,we find num_workers>0 can't work.

RuntimeError: DataLoader worker (pid 58353) exited unexpectedly with exit code 1. Details are lost due to multiprocessing. Rerunning with num_workers=0 may give better error trace.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the torchvision.io.decode_image call under DDP with device=local_rank and DataLoader num_workers>0. Rerun with num_workers=0 to obtain the detailed traceback, then compare the worker behavior. Done means identifying and resolving the worker failure while preserving multi-worker DDP loading.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
computer-vision, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.