THUDM / THUDM/slime

CUDA_ERROR_OUT_OF_MEMORY

Open
#1,497 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

为什么单独[Rank 3]的显存占用这么高?这是什么过程里的问题,要怎么解决呢?

Rollout generation: 100%|██████████| 256/256 [29:15<00:00,  1.52s/it]
Rollout generation: 100%|██████████| 256/256 [29:15<00:00,  6.86s/it]
(RolloutManager pid=516378) [2026-01-27 06:50:01] sglang_rollout.py:402 - Finish rollout: ['略'], label: 30, reward: 0.05
(RolloutManager pid=516378) [2026-01-27 06:50:01] sglang_rollout.py:311 - Abort request for http://33.184.121.56:15000
(SGLangEngine pid=516714) [2026-01-27 06:50:01] INFO:     33.184.121.56:52462 - "POST /abort_request HTTP/1.1" 200 OK
(RolloutManager pid=516378) [2026-01-27 06:50:01] rollout.py:534 - perf 13: {'perf/rollout_time': 1756.2989201545715, 'perf/tokens_per_gpu_per_sec': 35.01197848176559, 'perf/longest_sample_tokens_per_sec': 9.140243620139092, 'rollout/response_len/mean': 1921.609375, 'rollout/response_len/median': 784.0, 'rollout/zero_std/count_1.1': 4, 'rollout/zero_std/count_0.1': 1, 'rollout/repetition_frac': 0.0, 'rollout/truncated_ratio': 0.04296875}
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP4] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP7] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP3] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP6] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP2] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP0] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP5] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP1] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01] INFO:     33.184.121.56:52476 - "GET /flush_cache HTTP/1.1" 200 OK
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP6] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP1] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP7] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP0] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP3] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP4] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP2] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:01 TP5] Cache flushed successfully!
(SGLangEngine pid=516714) [2026-01-27 06:50:02] INFO:     33.184.121.56:52486 - "POST /release_memory_occupation HTTP/1.1" 200 OK
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:02] timer.py:24 - Timer wake_up start
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:02] memory_utils.py:41 - [Rank 0] Memory-Usage before wake_up model: {'gpu': '0', 'total_GB': 95.0, 'free_GB': 86.69, 'used_GB': 8.31, 'allocated_GB': 23.84, 'reserved_GB': 24.45}
(MegatronTrainRayActor pid=518532) [2026-01-27 06:50:08] reloadable_process_group.py:152 - Reloading 20 process groups in pid 518532
(MegatronTrainRayActor pid=518919) [2026-01-27 06:50:02] memory_utils.py:41 - [Rank 7] Memory-Usage before wake_up model: {'gpu': '7', 'total_GB': 95.0, 'free_GB': 88.55, 'used_GB': 6.46, 'allocated_GB': 24.33, 'reserved_GB': 25.32} [repeated 7x across cluster]
(MegatronTrainRayActor pid=518532) [2026-01-27 06:50:09] memory_utils.py:41 - [Rank 2] Memory-Usage after wake_up model: {'gpu': '2', 'total_GB': 95.0, 'free_GB': 63.55, 'used_GB': 31.45, 'allocated_GB': 24.42, 'reserved_GB': 25.38}
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:10] timer.py:32 - Timer wake_up end (elapsed: 8.0s)
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:10] timer.py:24 - Timer data_preprocess start
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:10] timer.py:32 - Timer data_preprocess end (elapsed: 0.1s)
(MegatronTrainRayActor pid=518321) AMEM pid:518321 sharedNetBuffersInit 561 dptr:0xda4ec00000 sz:268435456 defensive remove
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:12] timer.py:32 - Timer train_wait end (elapsed: 1803.2s)
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:12] timer.py:24 - Timer train start
(MegatronTrainRayActor pid=518321) [2026-01-27 06:50:12] timer.py:24 - Timer ref_log_probs start
(MegatronTrainRayActor pid=518321) AMEM pid:518321 sharedNetBuffersInit 561 dptr:0xdfa0000000 sz:268435456 defensive remove [repeated 16x across cluster]
(SGLangEngine pid=516714) [2026-01-27 06:50:56] INFO:     33.184.121.56:47738 - "GET /health HTTP/1.1" 200 OK
(MegatronTrainRayActor pid=518919) AMEM pid:518919 sharedNetBuffersInit 561 dptr:0xd6daa00000 sz:268435456 defensive remove [repeated 15x across cluster]
(MegatronTrainRayActor pid=518321) [2026-01-27 06:51:03] timer.py:32 - Timer ref_log_probs end (elapsed: 50.8s)
(MegatronTrainRayActor pid=518801) [2026-01-27 06:50:12] reloadable_process_group.py:152 - Reloading 20 process groups in pid 518801 [repeated 7x across cluster]
(MegatronTrainRayActor pid=518801) [2026-01-27 06:50:12] memory_utils.py:41 - [Rank 5] Memory-Usage after wake_up model: {'gpu': '5', 'total_GB': 95.0, 'free_GB': 63.54, 'used_GB': 31.47, 'allocated_GB': 24.32, 'reserved_GB': 25.39} [repeated 7x across cluster]
(MegatronTrainRayActor pid=518321) [2026-01-27 06:51:03] timer.py:24 - Timer log_probs start
(MegatronTrainRayActor pid=518321) [2026-01-27 06:51:21] timer.py:32 - Timer log_probs end (elapsed: 17.6s)
(MegatronTrainRayActor pid=518321) [2026-01-27 06:51:21] data.py:130 - rollout 13: {'rollout/raw_reward': 0.58203125, 'rollout/total_lengths': 3357.41015625, 'rollout/response_lengths': 2395.0390625, 'rollout/rewards': -0.0002327713300473988, 'rollout/truncated': 0.04296875, 'rollout/ref_log_probs': -0.3863580524921417, 'rollout/log_probs': -0.36543701589107513, 'rollout/advantages': 0.033923987532034516, 'rollout/returns': 0.033923987532034516}
(MegatronTrainRayActor pid=518321) [2026-01-27 06:51:21] timer.py:24 - Timer actor_train start
(MegatronTrainRayActor pid=518321) [torch_memory_saver.cpp] cuMemCreate CUDA_ERROR_OUT_OF_MEMORY (may not be an issue e.g. torch allocator will free cache and retry)
(SGLangEngine pid=516714) [2026-01-27 06:51:56] INFO:     33.184.121.56:38008 - "GET /health HTTP/1.1" 200 OK
(MegatronTrainRayActor pid=518672) [2026-01-27 06:52:37] memory_utils.py:41 - [Rank 3] Memory-Usage after torch distributed error: {'gpu': '3', 'total_GB': 95.0, 'free_GB': 0.01, 'used_GB': 95.0, 'allocated_GB': 23.95, 'reserved_GB': 82.69}

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the torch_memory_saver.cpp CUDA error and the memory snapshots from memory_utils.py:41, then trace the actor_train path around the reported Rank 3 failure. Compare the before/after wake_up and torch distributed error values across ranks; done means the cause of the asymmetric allocation is identified and a reproducible resolution is documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.