bigscience-workshop / bigscience-workshop/biomedical

Add S1000 corpus

Open
#933 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
505
Forks
117
PR merge metrics
No merged PRs in 30d

Description

## Adding a Dataset
- **Name:** S1000
- **Description:** The S1000 corpus is a comprehensive manual reannotation and extension of the S800 corpus that allows highly accurate recognition of species names, both for machine learning and dictionary-based methods.
- **Task:** NER
- **Paper:** https://academic.oup.com/bioinformatics/article/39/6/btad369/7192170?login=false
- **Data:** [https://jensenlab.org/resources/s1000/](https://jensenlab.org/resources/s1000/)
- **License:** Creative Commons Attribution 4.0 International

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.