iree-org / iree-org/iree

[CUDA] cuMemPrefetchAsync Error

Open
#15,664 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug 🐞
Dominant language
C++
Stars
3.9k
Forks
1k
Avg merge
4d 16h
Merged PRs (30d)
47

Description

### What happened?

I want to compile and run an example of matrix multiplication.
It compiles successfully in version A (commit id :`5abc05fd23efb109a2bf0170f47fd73cd01e2dad`) , but I get an error when running it. The error message says it is a arguments parsing error. as follows :
```
iree/runtime/src/iree/hal/drivers/cuda/cuda_allocator.c:339: INTERNAL; CUDA driver error 'CUDA_ERROR_INVALID_VALUE' (1): invalid argument; parsing value '40x60x90xf32=1'
```
This error is not present when running it on a newer version(commit id : `65817331c95a48ac3f7d972667f4bd2c35b8f61e`).
I debugged and found that the error occurred during the initialization of the gpu memory `cuMemPrefetchAsync` .
However, I couldn't find the exact change that caused this error to be fixed.

Who can help me locate the position of this update? (According to the commit ID, it seems that the updates were made between April 17th and July 11th.)

### Steps to reproduce your issue

## Input Test Case:
```
module attributes {torch.debug_module_name = "matmul"} {
func.func @forward(%arg0: tensor<40x60x90xf32>) -> tensor<40x60x80xf32> {
%0 = "tosa.const"() {value = dense<2.0> : tensor<40x90x80xf32>} : () -> tensor<40x90x80xf32>
%1 = "tosa.matmul"(%arg0, %0) : (tensor<40x60x90xf32>, tensor<40x90x80xf32>) -> tensor<40x60x80xf32>
return %1 : tensor<40x60x80xf32>
}
}
```
## old version cmd :
```
./iree-compile test.mlir --iree-input-type=tosa --iree-hal-target-backends=cuda -o a.vmfb

./iree-run-module --module=a.vmfb --device=cuda --function=forward --input="40x60x90xf32=1"
```
## old version result :
```
iree/runtime/src/iree/hal/drivers/cuda/cuda_allocator.c:339: INTERNAL; CUDA driver error 'CUDA_ERROR_INVALID_VALUE' (1): invalid argument; parsing value '40x60x90xf32=1'
```
## new version cmd:
```
./iree-compile test.mlir --iree-hal-target-backends=cuda -o a.vmfb

iree-run-module --module=a.vmfb --device=cuda --function=forward --input="40x60x90xf32=1"
```
## new version result:
```
EXEC @forward
result[0]: hal.buffer_view
40x60x80xf32=[[180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180 180][180 180 180 180 180 18 .......
```

### What component(s) does this issue relate to?

Runtime

### Version information

old version commit id: 5abc05fd23efb109a2bf0170f47fd73cd01e2dad

new version commit id: 65817331c95a48ac3f7d972667f4bd2c35b8f61e

### Additional context

_No response_

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with iree/runtime/src/iree/hal/drivers/cuda/cuda_allocator.c around line 339 and reproduce the old and new commands using the two referenced commits. Compare the changes between those commits while tracing cuMemPrefetchAsync and input parsing. Done means identifying the change that explains why the old command rejects 40x60x90xf32=1 and the newer version runs it.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, cpp
Domain
backend
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.