[Bug] Segmentation fault during run of quantized internlm/internlm-xcomposer2-4khd-7b model
- Dominant language
- Python
- Stars
- 8.1k
- Forks
- 748
- Avg merge
- 6d 2h
- Merged PRs (30d)
- 54
Description
### Checklist
- [ ] 1. I have searched related issues but cannot get the expected help.
- [ ] 2. The bug has not been fixed in the latest version.
### Describe the bug
Hello, after successful quantization of internlm/internlm-xcomposer2-4khd-7b model, I was trying to run VLM Offline Inference Pipeline. However, after loading of the model and the picture, it failed on segmentation fault. Do you have any idea, how to fix this issue?
### Reproduction
load_quant.py
```
import os
import torch
import nest_asyncio
from lmdeploy import pipeline, TurbomindEngineConfig
import logging
import requests
import cv2
import numpy as np
from io import BytesIO
# Enable logging
logging.basicConfig(level=logging.DEBUG)
nest_asyncio.apply()
# Set environment variables for debugging NCCL
os.environ["CUDA_VISIBLE_DEVICES"] = "0" # Run with a single GPU for debugging
os.environ['KMP_DUPLICATE_LIB_OK'] = 'True'
os.environ['NCCL_DEBUG'] = 'INFO'
os.environ['NCCL_DEBUG_SUBSYS'] = 'ALL'
os.environ['NCCL_P2P_LEVEL'] = 'NVL'
model_id = "internlm/internlm-xcomposer2-4khd-7b-chat-4bit"
backend_config = TurbomindEngineConfig(cache_max_entry_count=0.2, tp=1) # Set tp to 1 for single GPU
# Check CUDA and PyTorch installation
print(f"CUDA available: {torch.cuda.is_available()}")
print(f"Number of GPUs: {torch.cuda.device_count()}")
try:
pipe = pipeline(model_id, backend_config=backend_config)
print("Model loaded successfully")
except Exception as e:
print(f"Error loading model: {e}")
try:
image_url = 'https://raw.githubusercontent.com/open-mmlab/mmdeploy/main/tests/data/tiger.jpeg'
response = requests.get(image_url)
response.raise_for_status() # Raise an HTTPError for bad responses
image_array = np.array(bytearray(response.content), dtype=np.uint8)
img = cv2.imdecode(image_array, -1)
print(f"Image shape: {img.shape}, dtype: {img.dtype}")
except Exception as e:
print(f"Error downloading or processing image: {e}")
# Explicitly clear cache to manage memory
torch.cuda.empty_cache()
try:
prompts = [
{
'role': 'user',
'content': [
{'type': 'text', 'text': 'describe this image'},
{'type': 'image_url', 'image_url': {'url': image_url}}
]
}
]
response = pipe(prompts)
print(response)
except Exception as e:
print(f"Error processing image: {e}")
```
```
python3 load_quant.py
```
### Environment
```Shell
sys.platform: linux
Python: 3.10.12 (main, Mar 22 2024, 16:50:05) [GCC 11.4.0]
CUDA available: True
MUSA available: False
numpy_random_seed: 2147483648
GPU 0,1,2,3: NVIDIA A10G
CUDA_HOME: /usr/local/cuda-12.1
NVCC: Cuda compilation tools, release 12.1, V12.1.105
GCC: x86_64-linux-gnu-gcc (Ubuntu 11.4.0-1ubuntu1~22.04) 11.4.0
PyTorch: 2.2.2+cu121
PyTorch compiling details: PyTorch built with:
- GCC 9.3
- C++ Version: 201703
- Intel(R) oneAPI Math Kernel Library Version 2022.2-Product Build 20220804 for Intel(R) 64 architecture applications
- Intel(R) MKL-DNN v3.3.2 (Git Hash 2dc95a2ad0841e29db8b22fbccaf3e5da7992b01)
- OpenMP 201511 (a.k.a. OpenMP 4.5)
- LAPACK is enabled (usually provided by MKL)
- NNPACK is enabled
- CPU capability usage: AVX2
- CUDA Runtime 12.1
- NVCC architecture flags: -gencode;arch=compute_50,code=sm_50;-gencode;arch=compute_60,code=sm_60;-gencode;arch=compute_70,code=sm_70;-gencode;arch=compute_75,code=sm_75;-gencode;arch=compute_80,code=sm_80;-gencode;arch=compute_86,code=sm_86;-gencode;arch=compute_90,code=sm_90
- CuDNN 8.9.7 (built against CUDA 12.2)
- Built with CuDNN 8.9.2
- Magma 2.6.1
- Build settings: BLAS_INFO=mkl, BUILD_TYPE=Release, CUDA_VERSION=12.1, CUDNN_VERSION=8.9.2, CXX_COMPILER=/opt/rh/devtoolset-9/root/usr/bin/c++, CXX_FLAGS= -D_GLIBCXX_USE_CXX11_ABI=0 -fabi-version=11 -fvisibility-inlines-hidden -DUSE_PTHREADPOOL -DNDEBUG -DUSE_KINETO -DLIBKINETO_NOROCTRACER -DUSE_FBGEMM -DUSE_QNNPACK -DUSE_PYTORCH_QNNPACK -DUSE_XNNPACK -DSYMBOLICATE_MOBILE_DEBUG_HANDLE -O2 -fPIC -Wall -Wextra -Werror=return-type -Werror=non-virtual-dtor -Werror=bool-operation -Wnarrowing -Wno-missing-field-initializers -Wno-type-limits -Wno-array-bounds -Wno-unknown-pragmas -Wno-unused-parameter -Wno-unused-function -Wno-unused-result -Wno-strict-overflow -Wno-strict-aliasing -Wno-stringop-overflow -Wsuggest-override -Wno-psabi -Wno-error=pedantic -Wno-error=old-style-cast -Wno-missing-braces -fdiagnostics-color=always -faligned-new -Wno-unused-but-set-variable -Wno-maybe-uninitialized -fno-math-errno -fno-trapping-math -Werror=format -Wno-stringop-overflow, LAPACK_INFO=mkl, PERF_WITH_AVX=1, PERF_WITH_AVX2=1, PERF_WITH_AVX512=1, TORCH_VERSION=2.2.2, USE_CUDA=ON, USE_CUDNN=ON, USE_EXCEPTION_PTR=1, USE_GFLAGS=OFF, USE_GLOG=OFF, USE_MKL=ON, USE_MKLDNN=ON, USE_MPI=OFF, USE_NCCL=1, USE_NNPACK=ON, USE_OPENMP=ON, USE_ROCM=OFF, USE_ROCM_KERNEL_ASSERT=OFF,
TorchVision: 0.17.2+cu121
LMDeploy: 0.5.0+
transformers: 4.40.0
gradio: 3.50.2
fastapi: 0.111.0
pydantic: 2.7.4
triton: 2.2.0
```
### Error traceback
```Shell
DEBUG:asyncio:Using selector: EpollSelector
CUDA available: True
Number of GPUs: 1
Set max length to 16384
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): huggingface.co:443
DEBUG:urllib3.connectionpool:https://huggingface.co:443 "HEAD /openai/clip-vit-large-patch14-336/resolve/main/config.json HTTP/11" 200 0
DEBUG:urllib3.connectionpool:https://huggingface.co:443 "HEAD /openai/clip-vit-large-patch14-336/resolve/main/config.json HTTP/11" 200 0
Dummy Resized
INFO:accelerate.utils.modeling:We will use 90% of the memory on device 0 for storing the model, and 10% for the buffer to avoid OOM. You can set `max_memory` in to a higher value to use more memory (at your own risk).
[WARNING] gemm_config.in is not found; using default GEMM algo
Model loaded successfully
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): raw.githubusercontent.com:443
DEBUG:urllib3.connectionpool:https://raw.githubusercontent.com:443 "GET /open-mmlab/mmdeploy/main/tests/data/tiger.jpeg HTTP/11" 200 13929
Image shape: (182, 278, 3), dtype: uint8
DEBUG:urllib3.connectionpool:Starting new HTTPS connection (1): raw.githubusercontent.com:443
DEBUG:urllib3.connectionpool:https://raw.githubusercontent.com:443 "GET /open-mmlab/mmdeploy/main/tests/data/tiger.jpeg HTTP/11" 200 13929
Segmentation fault (core dumped)
```
Contributor guide
Assessment
This issue has not been assessed yet.