I had some problem dealing with multiple image inputs...
- Dominant language
- Python
- Stars
- 5.2k
- Forks
- 448
- Avg merge
- 3d 15h
- Merged PRs (30d)
- 26
Description
So here is the case, when I was trying to trun my dataset's format to the data format that more aligned with the xtuner, because I use mllm format, and I found xtuner use data format like this:
```
# well, from share-captioner_coco_lcs_sam_1246k_1107.json
{
"id": "sa_545504",
"image": "sam/images/sa_545504.jpg",
"conversations": [
{
"from": "human",
"value": "\nWhat is this photo about'?"
},
{
"from": "gpt",
"value": "The image captures a breathtaking view of the Hong Kong skyline at sunset. The sky, awash with hues of orange and blue, serves as a stunning backdrop to the city's architectural marvels. The tallest building, the International Finance Centre, pierces the sky with its imposing height. Its lights are switched on, casting a warm glow that contrasts with the cool tones of the evening sky.\n\nThe other buildings, though not as tall, are no less impressive. They are adorned with lights that twinkle like stars against the twilight sky. The water below mirrors the sky's colors, adding to the overall vibrancy of the scene.\n\nThe perspective of the image is from the water, looking towards the shore. This viewpoint allows for a comprehensive view of the cityscape, from the towering skyscrapers to the smaller structures nestled among them. The image encapsulates the essence of Hong Kong's urban landscape, a blend of modernity and natural beauty."
}
]
},
```
The "image" part only contains a string, which means there is only one image that can be included right? Can I do a list here?
I was trying to use the `llava_llama3_8b_instruct_clip_vit_large_p14_336_e1_gpu8_pretrain` template, but with my own dataset and images.
Contributor guide
Assessment
This issue has not been assessed yet.