DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3
Can i provide frame-index-level dialogue information when organizing my data?
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
hi! great job!
i want to finetune my videodata,but when i check the dataset prepare part, the annos are like this:
```
{
"video": ["videos/xxx.mp4"],
"conversations": [
{
"from": "human",
"value": "\nWhat are the main activities that take place in the video?"
},
{
"from": "gpt",
"value": "The main activities that take place in the video are the preparation of camera equipment by a man, a group of men riding a helicopter, and a man sailing a boat through the water."
},
...
]
},
```
these annos are video-level conversations,just tell us what the video says, but i want to know what happened in the video and also when does it happen.
so can i provide frame-index-level dialogue information when organizing my data?
For example: Does the person in the video pick up any object? If so, at what time segments does it happen?
Contributor guide
No contributing guide indexed for this repository
Research direction
Review the dataset prepare section and the video-level conversation schema shown in the issue. Determine whether frame-index or time-segment annotations are supported; done means establishing the accepted annotation format and the behavior needed for temporal questions.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100