PaddlePaddle / PaddlePaddle/FastDeploy

[Iluvatar] wint8 量化推理报错 CUINFER_STATUS_BAD_PARAM,且 debug 需要全量编译,耗时太长

Open
#7,063 8 comments 1 reaction 1 assignee View on GitHub

@zhupengyang is already working on this.

Since Mar 28, 2026.

Dominant language
Python
Stars
3.7k
Forks
756
Avg merge
19h 28m
Merged PRs (30d)
4

Description

问题描述

在 Iluvatar BI-V150S 上使用 FastDeploy 部署 ERNIE-4.5-21B-A3B-Paddle 模型,开启 --quantization wint8 后,服务启动失败,报错如下:

Error in file /home/aistudio/FastDeploy/custom_ops/iluvatar_ops/w8a16_group_gemv.cu on line 113: CUINFER_STATUS_BAD_PARAM
terminate called after throwing an instance of 'std::runtime_error'
what(): CUINFER_CHECK ERROR
环境信息
硬件: 4x Iluvatar BI-V150S (32GB)
操作系统: Ubuntu 20.04
FastDeploy: fastdeploy-iluvatar-gpu 2.5.0.dev0
PaddlePaddle: paddlepaddle 3.3.0
Paddle Iluvatar: paddle-iluvatar-gpu 3.3.0
aistudio-sdk: 0.3.8
模型: PaddlePaddle/ERNIE-4.5-21B-A3B-Paddle
启动命令
export PADDLE_XCCL_BACKEND=iluvatar_gpu
export INFERENCE_MSG_QUEUE_ID=232132
export LD_PRELOAD=/usr/local/corex/lib64/libcuda.so.1
export FD_SAMPLING_CLASS=rejection

python3 -m fastdeploy.entrypoints.openai.api_server \
       --model ./ERNIE-4.5-21B-A3B \
       --port 8180 \
       --tensor-parallel-size 4 \
       --quantization wint8 \
       --max-model-len 32768 \
       --block-size 16
调试困难

尝试通过修改源码添加 printf 来定位问题,但 FastDeploy 的编译方式只有 bash build.sh,每次修改都需要全量编译,耗时很长,调试效率极低。

期望
  1. 希望官方能够复现并修复这个 wint8 算子的 bug
  2. 如果可能,能否提供更轻量的编译方式(如增量编译)方便调试

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.