bytedance / bytedance/1d-tokenizer
visual representation quality of learned latent token
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 70
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
I appreciate such interesting work and really enjoyed your paper! Some questions have arisen while reading it:
1. I was impressed with the linear probing experiment on IN-1K (to my knowledge, it’s rarely seen in the image generation domain). I know this might be a bit off-topic, but have you tried, or do you have any insights on how it will works when fully fine-tuning the encoder and the latent token for image classification (just for measuring the quality of the learned representation)?
2. Did you train your own codebook?
Thanks! Hope this reaches you soon!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the linear probing experiment on IN-1K and the codebook training setup referenced in the issue. The issue names no files, tests, or entry points; done would require resolving the requested fine-tuning representation experiment and clarifying whether the codebook was trained by the project.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- jupyter-notebook
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100