bytedance / bytedance/1d-tokenizer

Training on larger datasets

Open
#30 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
70
PR merge metrics
No merged PRs in 30d

Description

Hi, congratulations on the great work, this approach looks very promising!
Are there any plans to release checkpoints trained on larger datasets?

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no files, tests, or training entry points. First clarify whether releasing larger-dataset checkpoints is planned and which datasets, training configuration, and artifacts are in scope. Done would require an agreed release plan and published checkpoints, rather than a narrowly defined code change.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.