AnswerDotAI / AnswerDotAI/ModernBERT

Training Tokenizer and pretraining

Open
#226 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.7k
Forks
145
PR merge metrics
No merged PRs in 30d

Description

Hi,

1. I'm looking to train a Portuguese version of modernBERT, likely both base and large models similar to the original. I noticed the paper mentions a modified OLMo tokenizer. What's the recommended approach for training a tokenizer for Portuguese in this context?

2. Regarding weight tiling for initializing the large model, is it enough to just follow the scripts provided in `bert_layers`?

3. Is there any documentation beyond the paper that covers the context extension and decay?

Thanks in advance!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reading the paper and the scripts in `bert_layers`, then identify the documented guidance needed for Portuguese tokenizer training, weight tiling for the large model, and context extension and decay. Done means providing documentation beyond the paper that answers all three questions; the issue has no named tests or additional entry points.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation, machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.