modelscope / modelscope/ms-swift

swift全参数预训练minicpm -v 2.6时,进程一直在同一步卡死,所有显卡的CPU占用均为100%

Open
#6,811 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
15.7k
Forks
1.7k
Avg merge
1d 16h
Merged PRs (30d)
136

Description

Describe the bug
swift全参数预训练minicpm -v 2.6时,进程一直在同一步卡死,所有显卡利用率均为100%

Image Image

Your hardware and system info
CUDA Version: 12.2
A800 8*40G显卡
torch==2.8.0

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No file, test, command, or traceback is provided. Start by reproducing MiniCPM-V 2.6 full-parameter pretraining with the reported CUDA 12.2, A800, and torch 2.8.0 setup, then collect the stuck-step logs and identify the responsible training entry point. Done means the hang is explained and the affected setup completes that step or has a verified workaround.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.