aymara / aymara/lima

Switch eng pos tagging learning corpus to OANC

Open
#16 0 comments 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
C++
Stars
119
Forks
20
PR merge metrics
No merged PRs in 30d

Description

The ANC MASC corpus is a lot larger than the NLTK WSJ subset that we currently use and it is really free making it easier to distribute.

We have to switch to this. This means mainly adapting it to lima tokenization (idioms and entities handled before learning the PoS tagging model).

Contributor guide

Open the contributing guide

Research direction

Start by locating the current English PoS-tagging training setup and the preprocessing that handles idioms and entities before learning. Replace the NLTK WSJ subset with the OANC/ANC MASC corpus and adapt it to lima tokenization; done means the model trains from the new corpus and the corpus can be distributed under its stated free terms.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
data, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.