Create Dataset readers for BC5CDR in BLUE Benchmark
- Dominant language
- Python
- Stars
- 253
- Forks
- 59
- PR merge metrics
- No merged PRs in 30d
Description
**Is your feature request related to a problem? Please describe.**
To prepare medical NER detection, we need to create a reader for the BC5CDR in the BLUE Benchmark: https://github.com/ncbi-nlp/BLUE_Benchmark
**Describe the solution you'd like**
1. Develop a reader for BC5CDR
2. Annotate the Entity Mentions from the dataset.
**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.
**Additional context**
Add any other context or screenshots about the feature request here.
Contributor guide
Research direction
Start by reviewing the BLUE Benchmark repository linked in the issue to understand the BC5CDR data format and required annotations. Implement a BC5CDR reader and annotate entity mentions from the dataset; done means the reader integrates with Forte's benchmark workflow and exposes the expected medical NER entities.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100