huggingface / huggingface/blog
Adding missing info to tutorial
- Dominant language
- Jupyter Notebook
- Stars
- 3.5k
- Forks
- 1k
- Avg merge
- 1d 20h
- Merged PRs (30d)
- 19
Description
This is a very helpful tutorial! Unfortunately there are a couple things missing that would help the clarity a lot.
1. The first would be to show what your config.json file is in ./models/EsperBERTo-small.
2. It's not entirely clear to me how to integrate EsperantoDataset into run_language_modeling.py. Showing the changes required to properly load the tokenizer would be very helpful.
Thanks again!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the tutorial section covering ./models/EsperBERTo-small and run_language_modeling.py, then inspect how EsperantoDataset is introduced. Document the expected config.json contents and the tokenizer-loading changes needed to integrate EsperantoDataset, using the tutorial's existing examples to verify the explanation is complete.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 38/100