duanyiqun / duanyiqun/Auto-ReID-Fast

Running with distributed

Open
#6 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
53
Forks
19
PR merge metrics
No merged PRs in 30d

Description

Hello, Duan.I am trying to repo this code .There are some questions when I use distributed. I can't use this command:
`srun -n your_node_nums --gres gpu:gpunums -p your_partition`
Error:
`The program 'srun' is currently not installed. To run 'srun' please ask your administrator to install the package 'slurm-client'
`
So when` CUDA out of memory`, what should I do to solve this problem. Looking forward to your help!

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue mentions the `srun` launch command and CUDA out-of-memory errors but names no files, tests, or entry points. Start by reading the distributed-running instructions and reproducing the reported command failure; the work is complete only when the project provides a clear, verified path for distributed execution and the reported memory problem.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.