[Bug] InternVL3-8B Quantization
- Dominant language
- Python
- Stars
- 8.1k
- Forks
- 748
- Avg merge
- 6d 2h
- Merged PRs (30d)
- 54
Description
### Checklist
- [ ] 1. I have searched related issues but cannot get the expected help.
- [ ] 2. The bug has not been fixed in the latest version.
- [ ] 3. Please note that if the bug-related issue you submitted lacks corresponding environment info and a minimal reproducible demo, it will be challenging for us to reproduce and resolve the issue, reducing the likelihood of receiving feedback.
### Describe the bug
internvl3-8b cannot be quantized on 3090.
### Reproduction
```
lmdeploy lite auto_awq \
OpenGVLab/InternVL3-8B \
--calib-samples 128 \
--calib-seqlen 2048 \
--w-bits 4 \
--w-group-size 128 \
--batch-size 1 \
--dtype bfloat16 \
--work-dir models/InternVL3-8B-lmdeploy-awq
```
### Environment
```Shell
pip list
Package Version
--------------------------------- ---------------------
accelerate 1.10.0
addict 2.4.0
aiohappyeyeballs 2.6.1
aiohttp 3.12.15
aiosignal 1.4.0
annotated-types 0.7.0
anyio 4.10.0
astor 0.8.1
async-timeout 5.0.1
attrs 25.3.0
blake3 1.0.5
cachetools 6.2.0
cbor2 5.7.0
certifi 2025.8.3
cffi 2.0.0
charset-normalizer 3.4.3
click 8.2.1
cloudpickle 3.1.1
compressed-tensors 0.11.1a20250912
cupy-cuda12x 13.6.0
datasets 4.1.0
depyf 0.19.0
dill 0.3.8
diskcache 5.6.3
distro 1.9.0
dnspython 2.8.0
einops 0.8.1
email-validator 2.3.0
exceptiongroup 1.3.0
fastapi 0.116.1
fastapi-cli 0.0.11
fastapi-cloud-cli 0.1.5
fastrlock 0.8.3
filelock 3.19.1
fire 0.7.1
flash_attn 2.8.2
flashinfer-python 0.3.1
frozendict 2.4.6
frozenlist 1.7.0
fsspec 2024.6.1
genson 1.3.0
gguf 0.17.1
h11 0.16.0
hf-xet 1.1.10
httpcore 1.0.9
httptools 0.6.4
httpx 0.28.1
huggingface-hub 0.35.0
idna 3.10
interegular 0.3.3
Jinja2 3.1.6
jiter 0.11.0
jsonpath-ng 1.7.0
jsonschema 4.25.1
jsonschema-specifications 2025.9.1
lark 1.2.2
llguidance 0.7.30
llmcompressor 0.7.2.dev39+g5e9b6d5e
llvmlite 0.44.0
lm-format-enforcer 0.11.3
lmdeploy 0.10.0
loguru 0.7.3
markdown-it-py 4.0.0
MarkupSafe 3.0.2
mdurl 0.1.2
mistral_common 1.8.5
mmengine-lite 0.10.7
mpmath 1.3.0
msgpack 1.1.1
msgspec 0.19.0
multidict 6.6.4
multiprocess 0.70.16
networkx 3.4.2
ninja 1.13.0
numba 0.61.2
numpy 1.26.4
nvidia-cublas-cu12 12.8.4.1
nvidia-cuda-cupti-cu12 12.8.90
nvidia-cuda-nvrtc-cu12 12.8.93
nvidia-cuda-runtime-cu12 12.8.90
nvidia-cudnn-cu12 9.10.2.21
nvidia-cudnn-frontend 1.14.1
nvidia-cufft-cu12 11.3.3.83
nvidia-cufile-cu12 1.13.1.3
nvidia-curand-cu12 10.3.9.90
nvidia-cusolver-cu12 11.7.3.90
nvidia-cusparse-cu12 12.5.8.93
nvidia-cusparselt-cu12 0.7.1
nvidia-ml-py 12.575.51
nvidia-nccl-cu12 2.27.3
nvidia-nvjitlink-cu12 12.8.93
nvidia-nvtx-cu12 12.8.90
openai 1.107.3
openai-harmony 0.0.4
opencv-python-headless 4.12.0.88
outlines 1.2.5
outlines_core 0.2.11
packaging 25.0
pandas 2.3.2
partial-json-parser 0.2.1.1.post6
peft 0.14.0
pillow 10.4.0
pip 25.2
platformdirs 4.4.0
ply 3.11
prometheus_client 0.22.1
prometheus-fastapi-instrumentator 7.1.0
propcache 0.3.2
protobuf 6.32.1
psutil 7.0.0
py-cpuinfo 9.0.0
pyarrow 21.0.0
pybase64 1.4.2
pycountry 24.6.1
pycparser 2.23
pydantic 2.11.9
pydantic_core 2.33.2
pydantic-extra-types 2.10.5
Pygments 2.19.2
pynvml 12.0.0
python-dateutil 2.9.0.post0
python-dotenv 1.1.1
python-json-logger 3.3.0
python-multipart 0.0.20
pytz 2025.2
PyYAML 6.0.2
pyzmq 27.1.0
ray 2.49.1
referencing 0.36.2
regex 2025.9.1
requests 2.32.5
rich 14.1.0
rich-toolkit 0.15.1
rignore 0.6.4
rpds-py 0.27.1
safetensors 0.6.2
scipy 1.15.3
sentencepiece 0.2.1
sentry-sdk 2.38.0
setproctitle 1.3.7
setuptools 80.9.0
shellingham 1.5.4
shortuuid 1.0.13
six 1.17.0
sniffio 1.3.1
soundfile 0.13.1
soxr 1.0.0
starlette 0.47.3
sympy 1.14.0
tabulate 0.9.0
termcolor 3.1.0
tiktoken 0.11.0
timm 1.0.19
tokenizers 0.21.4
tomli 2.2.1
torch 2.8.0
torchaudio 2.8.0
torchvision 0.23.0
tqdm 4.67.1
transformers 4.55.2
triton 3.4.0
typer 0.17.4
typing_extensions 4.15.0
typing-inspection 0.4.1
tzdata 2025.2
urllib3 2.5.0
uvicorn 0.35.0
uvloop 0.21.0
vllm 0.10.2
watchfiles 1.1.0
websockets 15.0.1
wheel 0.45.1
xformers 0.0.32.post1
xgrammar 0.1.23
xxhash 3.5.0
yapf 0.43.0
yarl 1.20.1
zstandard 0.25.0
```
### Error traceback
```Shell
Move model.embed_tokens to GPU.
Move model.layers.0 to CPU.
Move model.layers.1 to CPU.
Move model.layers.2 to CPU.
Move model.layers.3 to CPU.
Move model.layers.4 to CPU.
Move model.layers.5 to CPU.
Move model.layers.6 to CPU.
Move model.layers.7 to CPU.
Move model.layers.8 to CPU.
Move model.layers.9 to CPU.
Move model.layers.10 to CPU.
Move model.layers.11 to CPU.
Move model.layers.12 to CPU.
Move model.layers.13 to CPU.
Move model.layers.14 to CPU.
Move model.layers.15 to CPU.
Move model.layers.16 to CPU.
Move model.layers.17 to CPU.
Move model.layers.18 to CPU.
Move model.layers.19 to CPU.
Move model.layers.20 to CPU.
Move model.layers.21 to CPU.
Move model.layers.22 to CPU.
Move model.layers.23 to CPU.
Move model.layers.24 to CPU.
Move model.layers.25 to CPU.
Move model.layers.26 to CPU.
Move model.layers.27 to CPU.
Move model.norm to GPU.
Move model.rotary_emb to GPU.
Move lm_head to CPU.
Loading calibrate dataset ...
`trust_remote_code` is not supported anymore.
Please check that the Hugging Face dataset 'ptb_text_only' isn't based on a loading script and remove `trust_remote_code`.
If the dataset is based on a loading script, please ask the dataset author to remove it and convert it to a standard format like Parquet.
Using the latest cached version of the dataset since ptb_text_only couldn't be found on the Hugging Face Hub
Found the latest cached dataset configuration 'penn_treebank' at /root/.cache/huggingface/datasets/ptb_text_only/penn_treebank/1.1.0/8d1b97746fb9765d140e569ec5ddd35e20af4d37761f5e1bf357ea0b081f2c1f (last modified on Thu Sep 18 09:14:58 2025).
`trust_remote_code` is not supported anymore.
Please check that the Hugging Face dataset 'ptb_text_only' isn't based on a loading script and remove `trust_remote_code`.
If the dataset is based on a loading script, please ask the dataset author to remove it and convert it to a standard format like Parquet.
Using the latest cached version of the dataset since ptb_text_only couldn't be found on the Hugging Face Hub
Found the latest cached dataset configuration 'penn_treebank' at /root/.cache/huggingface/datasets/ptb_text_only/penn_treebank/1.1.0/8d1b97746fb9765d140e569ec5ddd35e20af4d37761f5e1bf357ea0b081f2c1f (last modified on Thu Sep 18 09:14:58 2025).
Token indices sequence length is longer than the specified maximum sequence length for this model (1104485 > 12288). Running this sequence through the model will result in indexing errors
model.layers.0, samples: 128, max gpu memory: 6.71 GB
Traceback (most recent call last):
File "/usr/local/bin/lmdeploy", line 7, in
sys.exit(run())
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/cli/entrypoint.py", line 39, in run
args.run(args)
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/cli/lite.py", line 111, in auto_awq
auto_awq(**kwargs)
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/lite/apis/auto_awq.py", line 86, in auto_awq
vl_model, model, tokenizer, work_dir = calibrate(model,
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/lite/apis/calibrate.py", line 318, in calibrate
calib_ctx.calibrate(all_data)
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/lite/quantization/calibration.py", line 224, in calibrate
_ = model(data.to(self.device))
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1773, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1784, in _call_impl
return forward_call(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/transformers/utils/generic.py", line 1083, in wrapper
outputs = func(self, *args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/transformers/models/qwen2/modeling_qwen2.py", line 379, in forward
hidden_states = decoder_layer(
File "/usr/local/lib/python3.10/dist-packages/transformers/modeling_layers.py", line 94, in __call__
return super().__call__(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1773, in _wrapped_call_impl
return self._call_impl(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/torch/nn/modules/module.py", line 1784, in _call_impl
return forward_call(*args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/lite/quantization/calibration.py", line 145, in _forward
batch_args, batch_kwargs = split_decoder_layer_inputs(self.batch_size, *args, **kwargs)
File "/usr/local/lib/python3.10/dist-packages/lmdeploy/lite/utils/batch_split.py", line 25, in split_decoder_layer_inputs
raise ValueError('The first argument must be a Tensor')
ValueError: The first argument must be a Tensor
```
Contributor guide
Assessment
This issue has not been assessed yet.