DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

Can i provide frame-index-level dialogue information when organizing my data?

Open
#93 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

hi! great job!
i want to finetune my videodata,but when i check the dataset prepare part, the annos are like this:
```
{
"video": ["videos/xxx.mp4"],
"conversations": [
{
"from": "human",
"value": "\nWhat are the main activities that take place in the video?"
},
{
"from": "gpt",
"value": "The main activities that take place in the video are the preparation of camera equipment by a man, a group of men riding a helicopter, and a man sailing a boat through the water."
},
...
]
},
```

these annos are video-level conversations,just tell us what the video says, but i want to know what happened in the video and also when does it happen.
so can i provide frame-index-level dialogue information when organizing my data?
For example: Does the person in the video pick up any object? If so, at what time segments does it happen?

Contributor guide

No contributing guide indexed for this repository

Research direction

Review the dataset prepare section and the video-level conversation schema shown in the issue. Determine whether frame-index or time-segment annotations are supported; done means establishing the accepted annotation format and the behavior needed for temporal questions.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.