allenai / allenai/S2AND

Future improvements

Open
#13 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
115
Forks
24
PR merge metrics
No merged PRs in 30d

Description

- [ ] Unify the set of languages between cld2 and fasttext (see `unify_lang` branch for a start)
- [x] Audit the list of name pairs (noticed (maria, mary), (kathleen, katherine))
- [ ] Generally improve language detection on titles (would require a whole model)
- [ ] if a person has two very disjoint "personas", they will end up as two clusters. Probably not resolvable, but putting here anyway
- [ ] somehow do better with low information papers (e.g. no abstract, venue, affiliation, references)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.