asyml / asyml/forte

Create Dataset readers for BC5CDR in BLUE Benchmark

Open
#240 0 comments 0 reactions 0 assignees View on GitHub
good first issue help wanted topic: data
Dominant language
Python
Stars
253
Forks
59
PR merge metrics
No merged PRs in 30d

Description

**Is your feature request related to a problem? Please describe.**
To prepare medical NER detection, we need to create a reader for the BC5CDR in the BLUE Benchmark: https://github.com/ncbi-nlp/BLUE_Benchmark

**Describe the solution you'd like**
1. Develop a reader for BC5CDR
2. Annotate the Entity Mentions from the dataset.

**Describe alternatives you've considered**
A clear and concise description of any alternative solutions or features you've considered.

**Additional context**
Add any other context or screenshots about the feature request here.

Contributor guide

Open the contributing guide

Research direction

Start by reviewing the BLUE Benchmark repository linked in the issue to understand the BC5CDR data format and required annotations. Implement a BC5CDR reader and annotate entity mentions from the dataset; done means the reader integrates with Forte's benchmark workflow and exposes the expected medical NER entities.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.