InternLM / InternLM/InternLM-XComposer
模型推理性能优化
Open
- Dominant language
- Python
- Stars
- 2.9k
- Forks
- 175
- PR merge metrics
- No merged PRs in 30d
Description
感谢博主开源~
最近试用了InternLM-XComposer-VL-7b模型,效果很棒,就是推理的速度有点慢,目前使用V100进行推理,显存26G,耗时10s/条,想问下模型有什么推荐的推理加速方法么,还望博主给点建议
另外还尝试了internlm/internlm-xcomposer-7b-4bit模型,相同机器环境,显存从26G降到20G,耗时翻倍20s/条,不知道是不是我哪里设置的不对,推理上变慢了很多,这是为什么呢
以下是我使用的环境:
机器:V100
torch:2.1
cuda:11.8
python:3.9
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.