ROCm / ROCm/hip

HIP graph replay reads stale device memory for externally-written buffer on gfx950 (MI355X) — causes MEMORY_VIOLATION in LLM decode

Open
#3,902 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
4.4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

Environment

  • GPU: AMD MI355X (gfx950)
  • ROCm: 7.2.x (rocm7.2.4)
  • PyTorch: 2.10.0+rocm7.2.4
  • Setup: 8-GPU tensor-parallel LLM inference (TP=8) using PyTorch torch.cuda.CUDAGraph for decode

Bug Summary

When a device buffer is written via H2D copy on the default stream immediately before graph.replay() — also on the same default stream — the in-graph kernel reads a stale value from the previous step instead of the freshly written value.

This is not a stream-ordering bug (the H2D copy and replay are on the same stream in program order), not a dangling-pointer bug (the captured pointer equals the live buffer address), and not a concurrent-write bug (the buffer is arange both before and after replay). The stale read is localized to the HIP graph replay path itself.


Symptom

Intermittent Memory access fault by GPU node-N MEMORY_VIOLATION during LLM decode (CUDAGraph replay path), non-deterministic, absent in eager mode. Triggered by a small integer buffer (cu_seqlens_q, int32, ~513 elements) that is written correctly each step but read stale inside the graph.

Debug-agent wave dump at fault:

wave_NNNN: pc=0x... <_convert_req_index_to_global_index_kernel> (stopped, reason: MEMORY_VIOLATION)
scalar registers:
    s3: 000003e3    ← qo_end = 0x3e3 = 995  (STALE cumulative value from prior prefill step)
    s2: 0000008b    ← token_id = 139  (loop variable, should be ≤ 256 for decode)
    ...
vector registers:
    v3: [0] 0000772b [1] 0000772b ...   ← derived offset ≈ 29 000+, far OOB in token_indices

The buffer should contain arange(0, 257) (decode: 1 token per request). Instead the graph kernel read qo_end = 995, the cumulative value left by the immediately preceding prefill step.

Source location: for token_id in range(qo_start, qo_end) — the stale qo_end drives token_id to ~995, causing an out-of-bounds load in token_indices[token_id].


Evidence: buffer is physically correct; stale is in the graph's read path

We performed three independent probes to localize the stale:

1. Before-replay D2D snapshot (no host sync)

Right before graph.replay(), a GPU→GPU copy (non_blocking=True, same default stream) snapshots cu_seqlens_q into a persistent buffer. A daemon thread later saves it — no synchronization on the hot path.

Result (8 ranks × 3000 steps, warmup + steady-state): buffer is always arange(0, 257), max = 256, zero cumulative rows. The physical device memory is correct before replay.

2. After-replay D2D snapshot (rolling ring, no host sync)

Same approach, taken after replay, rolling ring of 512 steps overwritten every 100 steps.

Result (8 ranks × 512 steps covering the crash step): buffer is always arange(0, 257), max = 256, zero cumulative rows — even at the exact steps preceding the crash.

3. In-graph kernel register dump (ROCm debug agent)

HSA_TOOLS_LIB=librocm-debug-agent.so.2, graph mode preserved (not eager), --save-code-objects. Crash reproduced with the unmodified kernel; wave dump captured at fault.

Result: s3 = 0x3e3 = 995 = stale qo_end from the immediately preceding prefill step (which legitimately wrote cumulative ~988 into the buffer before the current decode step overwrote it with arange). The buffer is provably arange (probes 1 + 2), yet the graph kernel read 995.

Conclusion from probes
Probe Buffer content Kernel read
Before-replay D2D arange (max 256)
After-replay D2D arange (max 256)
Debug-agent register 995 (stale cumulative)

Same address (ADDR_PROBE confirmed capture pointer == live pointer at every replay). Same physical memory (D2D probes confirm arange both sides of replay). Different value inside the graph. The stale is 100% in the graph replay's memory read path.


Characteristics

  • CUDAGraph required: eager mode runs 145 000+ decode steps with zero violations.
  • Multi-GPU TP=8 required: we attempted single-GPU minimal repros (torch copy graph, triton kernel graph, 500-kernel large graph, shared-pool multi-graph, all with external H2D write before replay) — none reproduced. The bug requires the real TP=8 setup with RCCL collectives inside the graph.
  • Non-deterministic / flaky: same commit alternates pass/fail; concurrent-scale-dependent (higher batch = more prefill-decode interleaving = higher stale probability).
  • One-epoch lag: the stale value is specifically the value written by the immediately preceding prefill step — not the capture-time value, not an arbitrary old value. This suggests the graph's view of the buffer lags by one write epoch.
  • In-kernel instrumentation masks the bug (Heisenbug): adding a tl.store to record the loaded value (and clamping the loop to avoid the crash) causes the bug to disappear — the recompiled/recaptured kernel stops exhibiting the stale. This makes in-kernel probes unreliable for this bug; the debug-agent register dump (which does not modify the kernel binary) is the only reliable in-kernel observable.
  • Graph pool: multiple graphs (one per decode batch size) share a single graph_pool_handle.

H2D write details

The write path:

# CpuGpuBuffer.copy_to_gpu()
self.gpu.copy_(self.cpu, non_blocking=True)   # pinned H2D on current stream

Called in prepare_inputs on the default stream, immediately before graph.replay() also on the default stream. Program-order synchronization should guarantee visibility. The forward pass is single-stream synchronous with no side-stream switching.


Ask

  1. Is this a known HIP-graph limitation on gfx950 (e.g., graph replay does not guarantee observability of H2D writes on the same stream made between replays)?
  2. Is there a timeline for a fix in HIP runtime or the graph replay path?
  3. Is there a recommended pattern to force the graph to observe external writes (other than inserting a cudaStreamSynchronize between the write and replay, which would defeat the performance purpose of CUDAGraph)?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No repository file or test is named. Start with the H2D path in CpuGpuBuffer.copy_to_gpu(), its prepare_inputs caller, and graph.replay() on the default stream; then reproduce the TP=8 gfx950 case while tracing HIP graph replay visibility. Done means the limitation or regression is localized and a supported fix or synchronization pattern is identified.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
devtools, performance
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.