InternLM / InternLM/lmdeploy

[Bug] CUDA runtime error: an illegal memory access was encountered

Open
#696 3 comments 0 reactions 1 assignee Claimed by @lzhangzz View on GitHub
Dominant language
Python
Stars
8.1k
Forks
748
Avg merge
6d 2h
Merged PRs (30d)
54

Description

### Checklist

- [ ] 1. I have searched related issues but cannot get the expected help.
- [ ] 2. The bug has not been fixed in the latest version.

### Describe the bug

模型:llama2-70B
设备:A100/40G × 4
lmdeploy版本:0.0.13

allocator对象貌似存在bug,持续运行时,有两种情况抛出错误了:
1. LlamaTritonModelInstance对象析构时,调用allocator->free(),出现段错误。
![Snipaste_2023-11-16_10-41-28](https://github.com/InternLM/lmdeploy/assets/103549265/0953ceff-36dc-41d7-a416-139837e9999a)

3. 内部线程执行ContextDecode时,调用allocator->malloc(),出现cuda runtime error。
![image](https://github.com/InternLM/lmdeploy/assets/103549265/0ece8314-347e-40ae-9d27-8c824c0a4695)

以上错误都是运行过程中随机出现的,可以正常处理一些请求。

程序是在一台cuda11.7版本的机器上编译,移到另一台cuda11.3的机器上运行的,有可能是cuda版本不一致而引起的吗?

### Reproduction

```c++

void function() {
std::vector> model_instances;
std::vector cuda_streams;
std::vector threads;

//创建model_instances
model_instances.resize((size_t)gpu_count);
cuda_streams.resize((size_t)gpu_count);
threads.clear();
for (int device_id = 0; device_id < gpu_count; device_id++) {
const int rank = node_id * gpu_count + device_id;
threads.emplace_back([this, device_id, rank, &model_instances, &cuda_streams]() {
ft::check_cuda_error(cudaSetDevice(device_id));
cudaStream_t stream;
ft::check_cuda_error(cudaStreamCreate(&stream));
cuda_streams.at(device_id) = stream;

auto model_instance = this->model->createModelInstance(device_id, rank, stream, this->nccl_comms, nullptr);
model_instances.at(device_id) = std::move(model_instance);
printf("model instance %d is created \n", device_id);
ft::print_mem_usage();
});
}
for (auto& t : threads) {
t.join();
}

//构造请求

//推理
threads.clear();
for (int device_id = 0; device_id < gpu_count; device_id++) {
threads.push_back(std::thread(threadForward,
&model_instances[device_id],
request_list[device_id],
&output_tensors_lists[device_id],
device_id,
instance_comm.get(),
node_id,
(void*)(&lmDeployRequest)));
}
for (auto& t : threads) {
t.join();
}

//释放model_instances
model_instances.clear();

//销毁句柄
for(int device_id = 0; device_id < gpu_count; device_id++) {
ft::check_cuda_error(cudaSetDevice(device_id));
cudaStream_t stream = cuda_streams.at(device_id);
ft::check_cuda_error(cudaStreamDestroy(stream));
}

//释放请求
}
```
以上是调用AbstractTransformerModelInstance进行推理的大致方式,劳烦看看有没有问题。

### Environment

```Shell
ubuntu-16.04
cuda-11.4
```

### Error traceback

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.