The 4B-pretrained model can not recognize video input.
Open
- Dominant language
- Python
- Stars
- 726
- Forks
- 48
- PR merge metrics
- No merged PRs in 30d
Description
Thanks for your great work!
I download the 4B-pretrained and 8B-pretrained checkpoints, and convert them to huggingface format.
Then I feed a video into them, respectively. They are told to describe the video.
However, the 4B model always outputs "The image is a photograph...", while the 8B model correctly outputs "The video appears to be a still shot from a video..."
So, does the 4B model after pretraining can not recognize video input?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.