google-deepmind / google-deepmind/videoprism

Question: How to Use VideoPrism for Video Captioning?

Open
#45 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
392
Forks
39
Avg merge
14h 35m
Merged PRs (30d)
1

Description

Hi, thank you for the great work on this project!

I’ve already had a chance to try out VideoPrism for retrieving the most relevant video based on a text description, following the [example notebook](https://github.com/google-deepmind/videoprism/blob/main/videoprism/colabs/videoprism_video_text_demo.ipynb) — it worked very well.

On the [Google Research Blog](https://research.google/blog/videoprism-a-foundational-visual-encoder-for-video-understanding/), I also saw that VideoPrism can be used for video captioning. Could you please let me know if this functionality is currently available, and if so, how it can be used?

Thanks in advance!

Contributor guide

Open the contributing guide

Research direction

Start with videoprism/colabs/videoprism_video_text_demo.ipynb and compare its demonstrated retrieval workflow with the Google Research Blog's captioning claim. Confirm whether captioning is available in this repository and document the supported usage, or clarify that it is not currently provided.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
machine-learning
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.