google-deepmind / google-deepmind/videoprism
Question: How to Use VideoPrism for Video Captioning?
- Dominant language
- Python
- Stars
- 392
- Forks
- 39
- Avg merge
- 14h 35m
- Merged PRs (30d)
- 1
Description
Hi, thank you for the great work on this project!
I’ve already had a chance to try out VideoPrism for retrieving the most relevant video based on a text description, following the [example notebook](https://github.com/google-deepmind/videoprism/blob/main/videoprism/colabs/videoprism_video_text_demo.ipynb) — it worked very well.
On the [Google Research Blog](https://research.google/blog/videoprism-a-foundational-visual-encoder-for-video-understanding/), I also saw that VideoPrism can be used for video captioning. Could you please let me know if this functionality is currently available, and if so, how it can be used?
Thanks in advance!
Contributor guide
Research direction
Start with videoprism/colabs/videoprism_video_text_demo.ipynb and compare its demonstrated retrieval workflow with the Google Research Blog's captioning claim. Confirm whether captioning is available in this repository and document the supported usage, or clarify that it is not currently provided.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- jupyter-notebook, python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100