huggingface / huggingface/blog

Adding missing info to tutorial

Open
#48 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
3.5k
Forks
1k
Avg merge
1d 20h
Merged PRs (30d)
19

Description

This is a very helpful tutorial! Unfortunately there are a couple things missing that would help the clarity a lot.
1. The first would be to show what your config.json file is in ./models/EsperBERTo-small.
2. It's not entirely clear to me how to integrate EsperantoDataset into run_language_modeling.py. Showing the changes required to properly load the tokenizer would be very helpful.

Thanks again!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the tutorial section covering ./models/EsperBERTo-small and run_language_modeling.py, then inspect how EsperantoDataset is introduced. Document the expected config.json contents and the tokenizer-loading changes needed to integrate EsperantoDataset, using the tutorial's existing examples to verify the explanation is complete.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.