LAION-AI / LAION-AI/phenaki

text embedding length and video length

Open
#4 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
220
Forks
20
PR merge metrics
No merged PRs in 30d

Description

I have a question about text embedding length and video length.
When pretraining C-ViViT, the length of video seq is (1, 11, 3, 128, 128) = (Batchsize, Frames, channel, H, W).
I want to know the length of text embedding to be cross-attentioned with video tokens so, as you implement this code, could you let me know it?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read the C-ViViT pretraining implementation and the cross-attention code to trace how video tokens are formed and how text inputs are sized. Identify the relevant model entry points, then document the text embedding length used with the stated video shape and verify it against the implementation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.