THUDM / THUDM/slime

[Bug] VLM multi-turn rollout: model cannot see images on multi-turns

Open
#1,847 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

Bug Description

(RolloutManager pid=276038) [2026-04-20 22:10:23] rollout.py:353 - [1][response_text]: I must have misjudged the situation. The image is blank — perhaps it's a placeholder or a test case to check if I can handle no input...
(RolloutManager pid=276038) Wait — perhaps I misread the actual image? The image was pasted as:
(RolloutManager pid=276038) The text given is only "The The The The The The The The The The ..." (repeated ~5000 times)

Steps to Reproduce
  1. 使用镜像安装
  2. Run the geo3k_vlm_multi_turn example with a VLM model (e.g., Qwen3-VL-4B-Instruct):
    python -m examples.geo3k_vlm_multi_turn.run_geo3k_vlm_multi_turn
    
    
Expected Behavior

看到图

Actual Behavior

看不到图

Environment
  • slime version: latest (commit 0988f0f4)
  • Model: Qwen3-VL-4B-Instruct
  • GPU: 8x GPUs A100
  • SGLang backend with sglang-router
  • Training backend: Megatron

rollout_log.txt

Logs

Additional Context

root@dsw-762422-7c87cbbcb9-84qp4:~/slime# pip list
Package Version Editable project location


absl-py 2.4.0
accelerate 1.12.0
addict 2.4.0
aiohappyeyeballs 2.6.1
aiohttp 3.13.3
aiohttp-cors 0.8.1
aiosignal 1.4.0
airportsdata 20260208
annotated-doc 0.0.4
annotated-types 0.7.0
anthropic 0.83.0
antlr4-python3-runtime 4.9.3
anyio 4.13.0
apache-tvm-ffi 0.1.8.post2
apex 0.1
argcomplete 3.6.3
asttokens 3.0.1
attrs 26.1.0
av 16.1.0
black 26.1.0
blinker 1.7.0
blobfile 3.0.0
build 1.4.0
cache_dit 1.2.0
certifi 2026.1.4
cffi 2.0.0
cfgv 3.5.0
charset-normalizer 3.4.4
click 8.3.1
cloudpickle 3.1.2
colorful 0.5.8
comm 0.2.3
compressed-tensors 0.13.0
contourpy 1.3.3
cryptography 41.0.7
cubloaty 0.1.0b3
cuda-bindings 12.9.5
cuda-pathfinder 1.3.4
cuda-python 12.9.0
cycler 0.12.1
datamodel-code-generator 0.54.0
datasets 4.5.0
dbus-python 1.3.2
debugpy 1.8.20
decorator 5.2.1
decord2 3.0.0
deep_ep 1.2.1
defusedxml 0.7.1
devscripts 2.23.7+ubuntu0.1
diffusers 0.36.0
dill 0.4.0
diskcache 5.6.3
distlib 0.4.0
distro 1.9.0
docstring_parser 0.17.0
einops 0.8.2
executing 2.2.1
fake_int4_quant_cuda 0.0.0
fastapi 0.135.2
filelock 3.24.3
fla-core 0.4.1
flash_attn 2.7.4.post1
flash_attn_3 3.0.0b1
flash-linear-attention 0.4.1
flashinfer-cubin 0.6.3
flashinfer-jit-cache 0.6.3+cu129
flashinfer-python 0.6.3
fonttools 4.61.1
frozenlist 1.8.0
fsspec 2025.10.0
genson 1.3.0
gguf 0.17.1
gitdb 4.0.12
GitPython 3.1.46
google-api-core 2.30.0
google-auth 2.48.0
google-cloud-core 2.5.0
google-cloud-storage 3.9.0
google-crc32c 1.8.0
google-resumable-media 2.8.0
googleapis-common-protos 1.72.0
grpcio 1.78.1
grpcio-health-checking 1.78.1
grpcio-reflection 1.78.1
h11 0.16.0
h2 4.3.0
hf_transfer 0.1.9
hf-xet 1.2.0
hpack 4.1.0
html5lib 1.1
httpcore 1.0.9
httplib2 0.20.4
httpx 0.28.1
httpx-sse 0.4.3
huggingface_hub 0.36.2
humanize 4.15.0
hydra-core 1.3.2
hyperframe 6.1.0
icdiff 2.0.10
identify 2.6.16
idna 3.11
imageio 2.36.0
imageio-ffmpeg 0.5.1
importlib_metadata 8.7.1
inflect 7.5.0
iniconfig 2.3.0
interegular 0.3.3
ipykernel 7.2.0
ipython 9.10.0
ipython_pygments_lexers 1.1.1
isort 7.0.0
jedi 0.19.2
Jinja2 3.1.6
jiter 0.13.0
jsonschema 4.26.0
jsonschema-specifications 2025.9.1
jupyter_client 8.8.0
jupyter_core 5.9.1
kiwisolver 1.4.9
lark 1.3.1
launchpadlib 1.11.0
lazr.restfulclient 0.14.6
lazr.uri 1.0.6
linkify-it-py 2.0.3
llguidance 0.7.30
llvmlite 0.46.0
loguru 0.7.3
lxml 6.0.2
Markdown 3.10.2
markdown-it-py 4.0.0
MarkupSafe 3.0.3
matplotlib 3.10.8
matplotlib-inline 0.2.1
maturin 1.12.4
mbridge 0.15.1
mcp 1.26.0
mdit-py-plugins 0.5.0
mdurl 0.1.2
megatron-bridge 0.3.0rc0
megatron-core 0.16.0rc0 /root/Megatron-LM
memray 1.19.1
ml_dtypes 0.5.4
modelscope 1.34.0
mooncake-transfer-engine 0.3.9
more-itertools 10.8.0
moviepy 2.2.1
mpmath 1.3.0
msgpack 1.1.2
msgspec 0.20.0
multidict 6.7.1
multiprocess 0.70.18
mypy_extensions 1.1.0
nest-asyncio 1.6.0
networkx 3.6.1
ninja 1.13.0
nixl 0.10.0
nixl-cu12 0.10.0
nodeenv 1.10.0
numba 0.64.0
numpy 1.26.4
nv-one-logger-core 2.3.1
nv-one-logger-training-telemetry 2.3.1
nvidia-cublas-cu12 12.9.1.4
nvidia-cuda-cupti-cu12 12.9.79
nvidia-cuda-nvrtc-cu12 12.9.86
nvidia-cuda-runtime-cu12 12.9.79
nvidia-cudnn-cu12 9.16.0.29
nvidia-cudnn-frontend 1.18.0
nvidia-cufft-cu12 11.4.1.4
nvidia-cufile-cu12 1.14.1.1
nvidia-curand-cu12 10.3.10.19
nvidia-cusolver-cu12 11.7.5.82
nvidia-cusparse-cu12 12.5.10.65
nvidia-cusparselt-cu12 0.7.1
nvidia-cutlass-dsl 4.3.5
nvidia-ml-py 13.590.48
nvidia-modelopt 0.41.0
nvidia-nccl-cu12 2.27.5
nvidia-nvjitlink-cu12 12.9.86
nvidia-nvshmem-cu12 3.3.20
nvidia-nvtx-cu12 12.9.79
nvidia-resiliency-ext 0.5.0
oauthlib 3.2.2
omegaconf 2.3.0
onnx 1.20.1
onnx-ir 0.2.0
onnxscript 0.6.2
openai 2.6.1
openai-harmony 0.0.4
opencensus 0.11.4
opencensus-context 0.1.3
opencv-python-headless 4.10.0.84
opentelemetry-api 1.39.1
opentelemetry-exporter-otlp 1.39.1
opentelemetry-exporter-otlp-proto-common 1.39.1
opentelemetry-exporter-otlp-proto-grpc 1.39.1
opentelemetry-exporter-otlp-proto-http 1.39.1
opentelemetry-exporter-prometheus 0.60b1
opentelemetry-proto 1.39.1
opentelemetry-sdk 1.39.1
opentelemetry-semantic-conventions 0.60b1
orjson 3.11.7
outlines 0.1.11
outlines_core 0.1.26
overrides 7.7.0
packaging 26.0
pandas 3.0.1
parso 0.8.6
partial-json-parser 0.2.1.1.post7
pathspec 1.0.4
pexpect 4.9.0
pillow 11.3.0
pip 26.0.1
platformdirs 4.9.2
pluggy 1.6.0
pre_commit 4.5.1
proglog 0.1.12
prometheus_client 0.24.1
prompt_toolkit 3.0.52
propcache 0.4.1
proto-plus 1.27.1
protobuf 6.33.5
psutil 7.2.2
ptyprocess 0.7.0
PuLP 3.3.0
pure_eval 0.2.3
py-spy 0.4.1
pyarrow 23.0.1
pyasn1 0.6.2
pyasn1_modules 0.4.2
pybase64 1.4.3
pycountry 26.2.16
pycparser 3.0
pycryptodomex 3.23.0
pydantic 2.12.5
pydantic_core 2.41.5
pydantic-settings 2.13.1
Pygments 2.19.2
PyGObject 3.48.2
PyJWT 2.11.0
pylatexenc 2.10
pyparsing 3.1.1
pyproject_hooks 1.2.0
pytest 9.0.2
python-apt 2.7.7+ubuntu5.2
python-dateutil 2.9.0.post0
python-dotenv 1.2.1
python-multipart 0.0.22
pytokens 0.4.1
PyYAML 6.0.3
pyzmq 27.1.0
quack-kernels 0.2.4
qwen-vl-utils 0.0.14
ray 2.54.0
referencing 0.37.0
regex 2026.2.19
remote-pdb 2.1.0
requests 2.32.5
rich 14.3.3
ring-flash-attn 0.1.8
rpds-py 0.30.0
rsa 4.9.1
runai-model-streamer 0.15.6
safetensors 0.7.0
scikit_build_core 0.11.6
scipy 1.17.1
sentencepiece 0.2.1
sentry-sdk 2.53.0
setproctitle 1.3.7
setuptools 82.0.0
sgl-kernel 0.3.21
sglang 0.5.9 /sgl-workspace/sglang/python
sglang-router 0.3.2
shellingham 1.5.4
six 1.17.0
slime 0.2.4 /mnt/cpfs_m6_29eu38p1/Group-m6/guantongkun.gtk/slime
smart_open 7.5.1
smg-grpc-proto 0.3.3
smmap 5.0.2
sniffio 1.3.1
soundfile 0.13.1
sse-starlette 3.2.0
st_attn 0.0.7
stack-data 0.6.3
starlette 1.0.0
StrEnum 0.4.15
sympy 1.14.0
tabulate 0.9.0
tensorboard 2.20.0
tensorboard-data-server 0.7.2
termplotlib 0.3.9
textual 8.0.0
tiktoken 0.12.0
tilelang 0.1.8
timm 1.0.16
tokenizers 0.22.2
toml 0.10.2
torch 2.9.1+cu129
torch_c_dlpack_ext 0.1.5
torch_memory_saver 0.0.9
torchao 0.9.0
torchaudio 2.9.1+cu129
torchcodec 0.8.0
torchvision 0.24.1+cu129
tornado 6.5.5
tqdm 4.67.3
traitlets 5.14.3
transformer_engine 2.10.0
transformer_engine_cu12 2.10.0
transformer_engine_torch 2.10.0
transformers 4.57.1
triton 3.5.1
typeguard 4.5.1
typer 0.24.1
typing_extensions 4.15.0
typing-inspection 0.4.2
uc-micro-py 1.0.3
urllib3 2.6.3
uv 0.10.4
uvicorn 0.42.0
uvloop 0.22.1
virtualenv 20.38.0
vsa 0.0.4
wadllib 1.3.6
wandb 0.25.0
wcwidth 0.6.0
webencodings 0.5.1
Werkzeug 3.1.6
wheel 0.46.3
wrapt 2.1.1
xgrammar 0.1.27
xxhash 3.6.0
yarl 1.23.0
z3-solver 4.15.4.0
zipp 3.23.0
root@dsw-762422-7c87cbbcb9-84qp4:~/slime#

Pre-submission Checklist
  • I have read the CONTRIBUTING.md and understand the collaboration scope.
  • I have read the documentation and my issue is not addressed there.
  • I have searched for existing issues and this is not a duplicate.
  • I have provided a minimal, reproducible example.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by running examples.geo3k_vlm_multi_turn.run_geo3k_vlm_multi_turn with the stated Qwen3-VL-4B-Instruct setup, then inspect rollout.py around line 353 and trace how images are handled across turns. Done means the model can see the image on multi-turn rollouts rather than receiving only repeated text.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.