DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

## Visual Grounding Support

Open
#76 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

@lixin4ever

I realized the annotation format is in below style. Does the architecture support visual grounding like Qwen-VL. If yes how do one prepare the jsonl annotation for training by adding bounding box annotations. :

```python
[
{
"image": ["images/xxx.jpg"],
"conversations": [
{
"from": "human",
"value": "\nWhat are the colors of the bus in the image?"
},
{
"from": "gpt",
"value": "The bus in the image is white and red."
},
...
]
},
{
"video": ["videos/xxx.mp4"],
"conversations": [
{
"from": "human",
"value": "\nWhat are the main activities that take place in the video?"
},
{
"from": "gpt",
"value": "The main activities that take place in the video are the preparation of camera equipment by a man, a group of men riding a helicopter, and a man sailing a boat through the water."
},
...
]
},
...
]
```

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.