bytedance / bytedance/1d-tokenizer
Experiments with video tokenization.
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 70
- PR merge metrics
- No merged PRs in 30d
Description
I made some changes to the model (3D convs) and trained the small one with 128 tokens on 128p 16-frame videos pre-compressed with CogvideoX's VAE and MSE loss.
Turned out better than I expected considering how fast the training was on consumer hardware (couple hours).
There's a lot of potential here, and I think I can improve the performance a lot further.





Contributor guide
No contributing guide indexed for this repository
Research direction
The issue describes experimental results but names no files, tests, entry points, or concrete acceptance criteria. Start by locating the model and training code for the reported 3D-convolution video-tokenization experiment; completion would require a defined follow-up objective, reproducible training setup, and agreed performance criteria.
Written by the indexing model from the issue text.
Assessment
- Domain
- computer-vision, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100