DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

video processing code failed

Open
#25 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

the code of processing video code is as
https://github.com/DAMO-NLP-SG/VideoLLaMA3/blob/22a53e317e057306adce199094bc9956904101de/videollama3/train.py#L248

It is inconsistent with the instructions to build dataset in readme file. In reademe file , the "video" is not a list but a string object.
```

[
{
"video": "images/xxx.jpg",
"conversations": [
{
"from": "human",
"value": "\nWhat are the colors of the bus in the image?"
},
{
"from": "gpt",
"value": "The bus in the image is white and red."
},
...
],
}
{
"video": "videos/xxx.mp4",
"conversations": [
{
"from": "human",
"value": "\nWhat are the main activities that take place in the video?"
},
{
"from": "gpt",
"value": "The main activities that take place in the video are the preparation of camera equipment by a man, a group of men riding a helicopter, and a man sailing a boat through the water."
},
...
],
},
...
]
```

Contributor guide

No contributing guide indexed for this repository

Research direction

Compare the dataset handling at train.py line 248 with the dataset format shown in the README, focusing on whether the video field is treated as a string or a list. Confirm the expected format for both image and video entries, then verify that the implementation and README describe the same structure.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.