aymara / aymara/lima

SpecificEntities with DeepNetwork models for other WikiNER languages

Open
#100 0 comments 0 reactions 0 assignees View on GitHub
enhancement help wanted wish
Dominant language
C++
Stars
119
Forks
20
PR merge metrics
No merged PRs in 30d

Description

The Deep pipeline for NamedEntity Recognition in Fre and Eng is based on NLP models trained on the [WikiNER datasets](https://figshare.com/articles/Learning_multilingual_named_entity_recognition_from_Wikipedia/5462500).

Since DeepLima povides tokenisation for more languages, we should use the other languages of WikiNER] to train new models :

- aij-wikiner-en . done
- aij-wikiner-fr . done
- aij-wikiner-de . todo
- aij-wikiner-es . todo
- aij-wikiner-it . todo
- aij-wikiner-nl . todo
- aij-wikiner-pl . todo
- aij-wikiner-pt . todo
- aij-wikiner-ru . todo

Contributor guide

Open the contributing guide

Research direction

Start by locating the Deep pipeline's named-entity-recognition model-training and model-integration entry points, then review the linked WikiNER dataset and the completed English and French models. Confirm how the German, Spanish, Italian, Dutch, Polish, Portuguese, and Russian models should be trained and registered; done means those listed todo languages have usable integrated models.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, machine-learning
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.