bytedance / bytedance/1d-tokenizer

Question about the gfid << rfid

Open
#46 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
70
PR merge metrics
No merged PRs in 30d

Description

Hi authors,

Thanks for sharing this interesting work.

I am curious about the relation between rfid and gfid presented by the recent works Maskbit and RAR. I noticed that the gfid can be significantly better compared to rfid (for example, rfig for RAR tokenizer is 2.28 while the gfid can be 1.48). Is there any explanation/discussion for this behavior? Thank you.

In addition, I noticed that the codebook size of RAR is set to 1024. Did you try to scale it up to a larger number?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue mentions Maskbit, RAR, rfid, gfid, and a 1024-entry codebook, but names no files or tests. Start by locating the reported tokenizer results and the implementation or configuration for the codebook size. Done would require a supported explanation of the metric difference and evidence from a larger-codebook experiment.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.