[Bug] VLM multi-turn rollout: model cannot see images on multi-turns
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.5k
- Forks
- 1.3k
- Avg merge
- 5h 36m
- Merged PRs (30d)
- 22
Description
Bug Description
(RolloutManager pid=276038) [2026-04-20 22:10:23] rollout.py:353 - [1][response_text]: I must have misjudged the situation. The image is blank — perhaps it's a placeholder or a test case to check if I can handle no input...
(RolloutManager pid=276038) Wait — perhaps I misread the actual image? The image was pasted as:
(RolloutManager pid=276038) The text given is only "The The The The The The The The The The ..." (repeated ~5000 times)
Steps to Reproduce
- 使用镜像安装
- Run the
geo3k_vlm_multi_turnexample with a VLM model (e.g., Qwen3-VL-4B-Instruct):python -m examples.geo3k_vlm_multi_turn.run_geo3k_vlm_multi_turn
Expected Behavior
看到图
Actual Behavior
看不到图
Environment
- slime version: latest (commit 0988f0f4)
- Model: Qwen3-VL-4B-Instruct
- GPU: 8x GPUs A100
- SGLang backend with sglang-router
- Training backend: Megatron
Logs
Additional Context
root@dsw-762422-7c87cbbcb9-84qp4:~/slime# pip list
Package Version Editable project location
absl-py 2.4.0
accelerate 1.12.0
addict 2.4.0
aiohappyeyeballs 2.6.1
aiohttp 3.13.3
aiohttp-cors 0.8.1
aiosignal 1.4.0
airportsdata 20260208
annotated-doc 0.0.4
annotated-types 0.7.0
anthropic 0.83.0
antlr4-python3-runtime 4.9.3
anyio 4.13.0
apache-tvm-ffi 0.1.8.post2
apex 0.1
argcomplete 3.6.3
asttokens 3.0.1
attrs 26.1.0
av 16.1.0
black 26.1.0
blinker 1.7.0
blobfile 3.0.0
build 1.4.0
cache_dit 1.2.0
certifi 2026.1.4
cffi 2.0.0
cfgv 3.5.0
charset-normalizer 3.4.4
click 8.3.1
cloudpickle 3.1.2
colorful 0.5.8
comm 0.2.3
compressed-tensors 0.13.0
contourpy 1.3.3
cryptography 41.0.7
cubloaty 0.1.0b3
cuda-bindings 12.9.5
cuda-pathfinder 1.3.4
cuda-python 12.9.0
cycler 0.12.1
datamodel-code-generator 0.54.0
datasets 4.5.0
dbus-python 1.3.2
debugpy 1.8.20
decorator 5.2.1
decord2 3.0.0
deep_ep 1.2.1
defusedxml 0.7.1
devscripts 2.23.7+ubuntu0.1
diffusers 0.36.0
dill 0.4.0
diskcache 5.6.3
distlib 0.4.0
distro 1.9.0
docstring_parser 0.17.0
einops 0.8.2
executing 2.2.1
fake_int4_quant_cuda 0.0.0
fastapi 0.135.2
filelock 3.24.3
fla-core 0.4.1
flash_attn 2.7.4.post1
flash_attn_3 3.0.0b1
flash-linear-attention 0.4.1
flashinfer-cubin 0.6.3
flashinfer-jit-cache 0.6.3+cu129
flashinfer-python 0.6.3
fonttools 4.61.1
frozenlist 1.8.0
fsspec 2025.10.0
genson 1.3.0
gguf 0.17.1
gitdb 4.0.12
GitPython 3.1.46
google-api-core 2.30.0
google-auth 2.48.0
google-cloud-core 2.5.0
google-cloud-storage 3.9.0
google-crc32c 1.8.0
google-resumable-media 2.8.0
googleapis-common-protos 1.72.0
grpcio 1.78.1
grpcio-health-checking 1.78.1
grpcio-reflection 1.78.1
h11 0.16.0
h2 4.3.0
hf_transfer 0.1.9
hf-xet 1.2.0
hpack 4.1.0
html5lib 1.1
httpcore 1.0.9
httplib2 0.20.4
httpx 0.28.1
httpx-sse 0.4.3
huggingface_hub 0.36.2
humanize 4.15.0
hydra-core 1.3.2
hyperframe 6.1.0
icdiff 2.0.10
identify 2.6.16
idna 3.11
imageio 2.36.0
imageio-ffmpeg 0.5.1
importlib_metadata 8.7.1
inflect 7.5.0
iniconfig 2.3.0
interegular 0.3.3
ipykernel 7.2.0
ipython 9.10.0
ipython_pygments_lexers 1.1.1
isort 7.0.0
jedi 0.19.2
Jinja2 3.1.6
jiter 0.13.0
jsonschema 4.26.0
jsonschema-specifications 2025.9.1
jupyter_client 8.8.0
jupyter_core 5.9.1
kiwisolver 1.4.9
lark 1.3.1
launchpadlib 1.11.0
lazr.restfulclient 0.14.6
lazr.uri 1.0.6
linkify-it-py 2.0.3
llguidance 0.7.30
llvmlite 0.46.0
loguru 0.7.3
lxml 6.0.2
Markdown 3.10.2
markdown-it-py 4.0.0
MarkupSafe 3.0.3
matplotlib 3.10.8
matplotlib-inline 0.2.1
maturin 1.12.4
mbridge 0.15.1
mcp 1.26.0
mdit-py-plugins 0.5.0
mdurl 0.1.2
megatron-bridge 0.3.0rc0
megatron-core 0.16.0rc0 /root/Megatron-LM
memray 1.19.1
ml_dtypes 0.5.4
modelscope 1.34.0
mooncake-transfer-engine 0.3.9
more-itertools 10.8.0
moviepy 2.2.1
mpmath 1.3.0
msgpack 1.1.2
msgspec 0.20.0
multidict 6.7.1
multiprocess 0.70.18
mypy_extensions 1.1.0
nest-asyncio 1.6.0
networkx 3.6.1
ninja 1.13.0
nixl 0.10.0
nixl-cu12 0.10.0
nodeenv 1.10.0
numba 0.64.0
numpy 1.26.4
nv-one-logger-core 2.3.1
nv-one-logger-training-telemetry 2.3.1
nvidia-cublas-cu12 12.9.1.4
nvidia-cuda-cupti-cu12 12.9.79
nvidia-cuda-nvrtc-cu12 12.9.86
nvidia-cuda-runtime-cu12 12.9.79
nvidia-cudnn-cu12 9.16.0.29
nvidia-cudnn-frontend 1.18.0
nvidia-cufft-cu12 11.4.1.4
nvidia-cufile-cu12 1.14.1.1
nvidia-curand-cu12 10.3.10.19
nvidia-cusolver-cu12 11.7.5.82
nvidia-cusparse-cu12 12.5.10.65
nvidia-cusparselt-cu12 0.7.1
nvidia-cutlass-dsl 4.3.5
nvidia-ml-py 13.590.48
nvidia-modelopt 0.41.0
nvidia-nccl-cu12 2.27.5
nvidia-nvjitlink-cu12 12.9.86
nvidia-nvshmem-cu12 3.3.20
nvidia-nvtx-cu12 12.9.79
nvidia-resiliency-ext 0.5.0
oauthlib 3.2.2
omegaconf 2.3.0
onnx 1.20.1
onnx-ir 0.2.0
onnxscript 0.6.2
openai 2.6.1
openai-harmony 0.0.4
opencensus 0.11.4
opencensus-context 0.1.3
opencv-python-headless 4.10.0.84
opentelemetry-api 1.39.1
opentelemetry-exporter-otlp 1.39.1
opentelemetry-exporter-otlp-proto-common 1.39.1
opentelemetry-exporter-otlp-proto-grpc 1.39.1
opentelemetry-exporter-otlp-proto-http 1.39.1
opentelemetry-exporter-prometheus 0.60b1
opentelemetry-proto 1.39.1
opentelemetry-sdk 1.39.1
opentelemetry-semantic-conventions 0.60b1
orjson 3.11.7
outlines 0.1.11
outlines_core 0.1.26
overrides 7.7.0
packaging 26.0
pandas 3.0.1
parso 0.8.6
partial-json-parser 0.2.1.1.post7
pathspec 1.0.4
pexpect 4.9.0
pillow 11.3.0
pip 26.0.1
platformdirs 4.9.2
pluggy 1.6.0
pre_commit 4.5.1
proglog 0.1.12
prometheus_client 0.24.1
prompt_toolkit 3.0.52
propcache 0.4.1
proto-plus 1.27.1
protobuf 6.33.5
psutil 7.2.2
ptyprocess 0.7.0
PuLP 3.3.0
pure_eval 0.2.3
py-spy 0.4.1
pyarrow 23.0.1
pyasn1 0.6.2
pyasn1_modules 0.4.2
pybase64 1.4.3
pycountry 26.2.16
pycparser 3.0
pycryptodomex 3.23.0
pydantic 2.12.5
pydantic_core 2.41.5
pydantic-settings 2.13.1
Pygments 2.19.2
PyGObject 3.48.2
PyJWT 2.11.0
pylatexenc 2.10
pyparsing 3.1.1
pyproject_hooks 1.2.0
pytest 9.0.2
python-apt 2.7.7+ubuntu5.2
python-dateutil 2.9.0.post0
python-dotenv 1.2.1
python-multipart 0.0.22
pytokens 0.4.1
PyYAML 6.0.3
pyzmq 27.1.0
quack-kernels 0.2.4
qwen-vl-utils 0.0.14
ray 2.54.0
referencing 0.37.0
regex 2026.2.19
remote-pdb 2.1.0
requests 2.32.5
rich 14.3.3
ring-flash-attn 0.1.8
rpds-py 0.30.0
rsa 4.9.1
runai-model-streamer 0.15.6
safetensors 0.7.0
scikit_build_core 0.11.6
scipy 1.17.1
sentencepiece 0.2.1
sentry-sdk 2.53.0
setproctitle 1.3.7
setuptools 82.0.0
sgl-kernel 0.3.21
sglang 0.5.9 /sgl-workspace/sglang/python
sglang-router 0.3.2
shellingham 1.5.4
six 1.17.0
slime 0.2.4 /mnt/cpfs_m6_29eu38p1/Group-m6/guantongkun.gtk/slime
smart_open 7.5.1
smg-grpc-proto 0.3.3
smmap 5.0.2
sniffio 1.3.1
soundfile 0.13.1
sse-starlette 3.2.0
st_attn 0.0.7
stack-data 0.6.3
starlette 1.0.0
StrEnum 0.4.15
sympy 1.14.0
tabulate 0.9.0
tensorboard 2.20.0
tensorboard-data-server 0.7.2
termplotlib 0.3.9
textual 8.0.0
tiktoken 0.12.0
tilelang 0.1.8
timm 1.0.16
tokenizers 0.22.2
toml 0.10.2
torch 2.9.1+cu129
torch_c_dlpack_ext 0.1.5
torch_memory_saver 0.0.9
torchao 0.9.0
torchaudio 2.9.1+cu129
torchcodec 0.8.0
torchvision 0.24.1+cu129
tornado 6.5.5
tqdm 4.67.3
traitlets 5.14.3
transformer_engine 2.10.0
transformer_engine_cu12 2.10.0
transformer_engine_torch 2.10.0
transformers 4.57.1
triton 3.5.1
typeguard 4.5.1
typer 0.24.1
typing_extensions 4.15.0
typing-inspection 0.4.2
uc-micro-py 1.0.3
urllib3 2.6.3
uv 0.10.4
uvicorn 0.42.0
uvloop 0.22.1
virtualenv 20.38.0
vsa 0.0.4
wadllib 1.3.6
wandb 0.25.0
wcwidth 0.6.0
webencodings 0.5.1
Werkzeug 3.1.6
wheel 0.46.3
wrapt 2.1.1
xgrammar 0.1.27
xxhash 3.6.0
yarl 1.23.0
z3-solver 4.15.4.0
zipp 3.23.0
root@dsw-762422-7c87cbbcb9-84qp4:~/slime#
Pre-submission Checklist
- I have read the CONTRIBUTING.md and understand the collaboration scope.
- I have read the documentation and my issue is not addressed there.
- I have searched for existing issues and this is not a duplicate.
- I have provided a minimal, reproducible example.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by running examples.geo3k_vlm_multi_turn.run_geo3k_vlm_multi_turn with the stated Qwen3-VL-4B-Instruct setup, then inspect rollout.py around line 353 and trace how images are handled across turns. Done means the model can see the image on multi-turn rollouts rather than receiving only repeated text.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100