bytedance / bytedance/1d-tokenizer
Question about the gfid << rfid
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 70
- PR merge metrics
- No merged PRs in 30d
Description
Hi authors,
Thanks for sharing this interesting work.
I am curious about the relation between rfid and gfid presented by the recent works Maskbit and RAR. I noticed that the gfid can be significantly better compared to rfid (for example, rfig for RAR tokenizer is 2.28 while the gfid can be 1.48). Is there any explanation/discussion for this behavior? Thank you.
In addition, I noticed that the codebook size of RAR is set to 1024. Did you try to scale it up to a larger number?
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue mentions Maskbit, RAR, rfid, gfid, and a 1024-entry codebook, but names no files or tests. Start by locating the reported tokenizer results and the implementation or configuration for the codebook size. Done would require a supported explanation of the metric difference and evidence from a larger-codebook experiment.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100