AnswerDotAI / AnswerDotAI/ModernBERT
Training Tokenizer and pretraining
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 145
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
1. I'm looking to train a Portuguese version of modernBERT, likely both base and large models similar to the original. I noticed the paper mentions a modified OLMo tokenizer. What's the recommended approach for training a tokenizer for Portuguese in this context?
2. Regarding weight tiling for initializing the large model, is it enough to just follow the scripts provided in `bert_layers`?
3. Is there any documentation beyond the paper that covers the context extension and decay?
Thanks in advance!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading the paper and the scripts in `bert_layers`, then identify the documented guidance needed for Portuguese tokenizer training, weight tiling for the large model, and context extension and decay. Done means providing documentation beyond the paper that answers all three questions; the issue has no named tests or additional entry points.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- documentation, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100