InternLM / InternLM/Tutorial

camp2 lmdeploy llava运行时输入高分辨率图片会返回空字符串

Open
#620 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
2k
Forks
1.5k
PR merge metrics
No merged PRs in 30d

Description

camp2 lmdeploy llava运行时输入高分辨率图片会返回空字符串

项目地址 https://github.com/InternLM/Tutorial/blob/camp2/lmdeploy/README.md#61-%E4%BD%BF%E7%94%A8lmdeploy%E8%BF%90%E8%A1%8C%E8%A7%86%E8%A7%89%E5%A4%9A%E6%A8%A1%E6%80%81%E5%A4%A7%E6%A8%A1%E5%9E%8Bllava

当输入一张分辨率为 1920*1080 分辨率的图片时不会返回文字
![llava4](https://github.com/InternLM/Tutorial/assets/72508155/7496d17e-8ac8-4281-b4d0-2cfd1f56255d)
打印response时显示text为空
![llava5](https://github.com/InternLM/Tutorial/assets/72508155/401f739c-5e77-4a36-862b-b9a1bddc1923)

解决方法为手动降低分辨率
```python
import gradio as gr
from lmdeploy import pipeline

# pipe = pipeline('liuhaotian/llava-v1.6-vicuna-7b') 非开发机运行此命令
pipe = pipeline('/share/new_models/liuhaotian/llava-v1.6-vicuna-7b')

def model(image, text):
if image is None:
return [(text, "请上传一张图片。")]
else:
width, height = image.size
print(f"width = {width}, height = {height}")

# 调整图片最长宽/高为256
if max(width, height) > 256:
ratio = max(width, height) / 256
n_width = int(width / ratio)
n_height = int(height / ratio)
print(f"new width = {n_width}, new height = {n_height}")
image = image.resize((n_width, n_height))

response = pipe((text, image)).text
print(f"response: {response}")
return [(text, response)]

demo = gr.Interface(fn=model, inputs=[gr.Image(type="pil"), gr.Textbox()], outputs=gr.Chatbot())
demo.launch()
```

更改后效果可以正常返回text
![llava6](https://github.com/InternLM/Tutorial/assets/72508155/1a4ccaba-b00e-41c5-91e1-0641e590e5a5)
![llava7](https://github.com/InternLM/Tutorial/assets/72508155/9a58bac6-1c20-426b-8387-2ba02accb0c5)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.