bytedance / bytedance/1d-tokenizer

Experiments with video tokenization.

Open
#37 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
70
PR merge metrics
No merged PRs in 30d

Description

I made some changes to the model (3D convs) and trained the small one with 128 tokens on 128p 16-frame videos pre-compressed with CogvideoX's VAE and MSE loss.
Turned out better than I expected considering how fast the training was on consumer hardware (couple hours).

There's a lot of potential here, and I think I can improve the performance a lot further.
![Untitled](https://github.com/user-attachments/assets/eb9c8f7a-0f90-4cee-8fcb-b3089b616e6f)
![Untitled-1](https://github.com/user-attachments/assets/cde9c6dd-cc72-4204-8a84-1580f06ecbd3)
![Untitled-2](https://github.com/user-attachments/assets/12f9c3d8-ca59-45ff-96f9-cd446796c95d)
![Untitled-5](https://github.com/user-attachments/assets/1309f40c-4cad-47c2-a47e-ba546a75df1c)
![Untitled-4](https://github.com/user-attachments/assets/987f8814-400d-42c3-9998-37179f788c77)

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue describes experimental results but names no files, tests, entry points, or concrete acceptance criteria. Start by locating the model and training code for the reported 3D-convolution video-tokenization experiment; completion would require a defined follow-up objective, reproducible training setup, and agreed performance criteria.

Written by the indexing model from the issue text.

Assessment

Domain
computer-vision, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.