Switch eng pos tagging learning corpus to OANC
Open
enhancement
- Dominant language
- C++
- Stars
- 119
- Forks
- 20
- PR merge metrics
- No merged PRs in 30d
Description
The ANC MASC corpus is a lot larger than the NLTK WSJ subset that we currently use and it is really free making it easier to distribute.
We have to switch to this. This means mainly adapting it to lima tokenization (idioms and entities handled before learning the PoS tagging model).
Contributor guide
Research direction
Start by locating the current English PoS-tagging training setup and the preprocessing that handles idioms and entities before learning. Replace the NLTK WSJ subset with the OANC/ANC MASC corpus and adapt it to lima tokenization; done means the model trains from the new corpus and the corpus can be distributed under its stated free terms.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100