CUDA_ERROR_OUT_OF_MEMORY
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.5k
- Forks
- 1.3k
- Avg merge
- 5h 36m
- Merged PRs (30d)
- 22
Description
为什么单独[Rank 3]的显存占用这么高?这是什么过程里的问题,要怎么解决呢?
Rollout generation: 100%|██████████| 256/256 [29:15<00:00, 1.52s/it]
Rollout generation: 100%|██████████| 256/256 [29:15<00:00, 6.86s/it]
[36m(RolloutManager pid=516378)[0m [2026-01-27 06:50:01] sglang_rollout.py:402 - Finish rollout: ['略'], label: 30, reward: 0.05
[36m(RolloutManager pid=516378)[0m [2026-01-27 06:50:01] sglang_rollout.py:311 - Abort request for http://33.184.121.56:15000
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01] INFO: 33.184.121.56:52462 - "POST /abort_request HTTP/1.1" 200 OK
[36m(RolloutManager pid=516378)[0m [2026-01-27 06:50:01] rollout.py:534 - perf 13: {'perf/rollout_time': 1756.2989201545715, 'perf/tokens_per_gpu_per_sec': 35.01197848176559, 'perf/longest_sample_tokens_per_sec': 9.140243620139092, 'rollout/response_len/mean': 1921.609375, 'rollout/response_len/median': 784.0, 'rollout/zero_std/count_1.1': 4, 'rollout/zero_std/count_0.1': 1, 'rollout/repetition_frac': 0.0, 'rollout/truncated_ratio': 0.04296875}
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP4] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP7] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP3] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP6] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP2] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP0] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP5] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP1] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01] INFO: 33.184.121.56:52476 - "GET /flush_cache HTTP/1.1" 200 OK
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP6] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP1] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP7] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP0] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP3] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP4] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP2] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:01 TP5] Cache flushed successfully!
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:02] INFO: 33.184.121.56:52486 - "POST /release_memory_occupation HTTP/1.1" 200 OK
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:02] timer.py:24 - Timer wake_up start
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:02] memory_utils.py:41 - [Rank 0] Memory-Usage before wake_up model: {'gpu': '0', 'total_GB': 95.0, 'free_GB': 86.69, 'used_GB': 8.31, 'allocated_GB': 23.84, 'reserved_GB': 24.45}
[36m(MegatronTrainRayActor pid=518532)[0m [2026-01-27 06:50:08] reloadable_process_group.py:152 - Reloading 20 process groups in pid 518532
[36m(MegatronTrainRayActor pid=518919)[0m [2026-01-27 06:50:02] memory_utils.py:41 - [Rank 7] Memory-Usage before wake_up model: {'gpu': '7', 'total_GB': 95.0, 'free_GB': 88.55, 'used_GB': 6.46, 'allocated_GB': 24.33, 'reserved_GB': 25.32}[32m [repeated 7x across cluster][0m
[36m(MegatronTrainRayActor pid=518532)[0m [2026-01-27 06:50:09] memory_utils.py:41 - [Rank 2] Memory-Usage after wake_up model: {'gpu': '2', 'total_GB': 95.0, 'free_GB': 63.55, 'used_GB': 31.45, 'allocated_GB': 24.42, 'reserved_GB': 25.38}
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:10] timer.py:32 - Timer wake_up end (elapsed: 8.0s)
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:10] timer.py:24 - Timer data_preprocess start
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:10] timer.py:32 - Timer data_preprocess end (elapsed: 0.1s)
[36m(MegatronTrainRayActor pid=518321)[0m AMEM pid:518321 sharedNetBuffersInit 561 dptr:0xda4ec00000 sz:268435456 defensive remove
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:12] timer.py:32 - Timer train_wait end (elapsed: 1803.2s)
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:12] timer.py:24 - Timer train start
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:50:12] timer.py:24 - Timer ref_log_probs start
[36m(MegatronTrainRayActor pid=518321)[0m AMEM pid:518321 sharedNetBuffersInit 561 dptr:0xdfa0000000 sz:268435456 defensive remove[32m [repeated 16x across cluster][0m
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:50:56] INFO: 33.184.121.56:47738 - "GET /health HTTP/1.1" 200 OK
[36m(MegatronTrainRayActor pid=518919)[0m AMEM pid:518919 sharedNetBuffersInit 561 dptr:0xd6daa00000 sz:268435456 defensive remove[32m [repeated 15x across cluster][0m
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:51:03] timer.py:32 - Timer ref_log_probs end (elapsed: 50.8s)
[36m(MegatronTrainRayActor pid=518801)[0m [2026-01-27 06:50:12] reloadable_process_group.py:152 - Reloading 20 process groups in pid 518801[32m [repeated 7x across cluster][0m
[36m(MegatronTrainRayActor pid=518801)[0m [2026-01-27 06:50:12] memory_utils.py:41 - [Rank 5] Memory-Usage after wake_up model: {'gpu': '5', 'total_GB': 95.0, 'free_GB': 63.54, 'used_GB': 31.47, 'allocated_GB': 24.32, 'reserved_GB': 25.39}[32m [repeated 7x across cluster][0m
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:51:03] timer.py:24 - Timer log_probs start
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:51:21] timer.py:32 - Timer log_probs end (elapsed: 17.6s)
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:51:21] data.py:130 - rollout 13: {'rollout/raw_reward': 0.58203125, 'rollout/total_lengths': 3357.41015625, 'rollout/response_lengths': 2395.0390625, 'rollout/rewards': -0.0002327713300473988, 'rollout/truncated': 0.04296875, 'rollout/ref_log_probs': -0.3863580524921417, 'rollout/log_probs': -0.36543701589107513, 'rollout/advantages': 0.033923987532034516, 'rollout/returns': 0.033923987532034516}
[36m(MegatronTrainRayActor pid=518321)[0m [2026-01-27 06:51:21] timer.py:24 - Timer actor_train start
[36m(MegatronTrainRayActor pid=518321)[0m [torch_memory_saver.cpp] cuMemCreate CUDA_ERROR_OUT_OF_MEMORY (may not be an issue e.g. torch allocator will free cache and retry)
[36m(SGLangEngine pid=516714)[0m [2026-01-27 06:51:56] INFO: 33.184.121.56:38008 - "GET /health HTTP/1.1" 200 OK
[36m(MegatronTrainRayActor pid=518672)[0m [2026-01-27 06:52:37] memory_utils.py:41 - [Rank 3] Memory-Usage after torch distributed error: {'gpu': '3', 'total_GB': 95.0, 'free_GB': 0.01, 'used_GB': 95.0, 'allocated_GB': 23.95, 'reserved_GB': 82.69}
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the torch_memory_saver.cpp CUDA error and the memory snapshots from memory_utils.py:41, then trace the actor_train path around the reported Rank 3 failure. Compare the before/after wake_up and torch distributed error values across ranks; done means the cause of the asymmetric allocation is identified and a reproducible resolution is documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100