Ramp up training for languages in the NLLB-200
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39
- Forks
- 7
- Avg merge
- 1d 9h
- Merged PRs (30d)
- 5
Description
There are some translation projects where both languages are in the NLLB-200. It may not be the best choice to take 3 drafted chapters of Mark and train the model for 20,000 steps. In that light, we should research a good path for this. Possibly this could involve investigating:
- Train on less steps (validation split)
- Use subset of FLORES-200 data with the new book (drag and drop into SF)
- Use different training weights for the different text (in the SF UI?)
- Add a tag for Scripture and a tag for "other stuff" (drop-down in the SF UI?)
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by examining the current NLLB-200 training flow and the SF UI options for validation splits, FLORES-200 subsets, training weights, and Scripture/other tags. Compare these alternatives for projects where both languages are in NLLB-200. Done means documenting a recommended path for using drafted chapters with the model.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100