aboutcode-org / aboutcode-org/scancode-toolkit
[GSoC] Build NER training dataset from annotated rules
Đang mở
GSoC
- Ngôn ngữ chính
- Python
- Star
- 2.6k
- Fork
- 791
- Merge trung bình
- 1 ngày 12 giờ
- Pull request đã merge (30 ngày)
- 5
Mô tả
Tasks:
- Parse .RULE files and extract extract required phrase markers
- Propagate annotations (with CLI commands from #3924 ) to unannotated rules (with matching text) to expand the dataset
- Improve composite rules (AND/OR expressions where some aren't marked)
- Convert to BIOES labels
- Split into train/val/test by rule family
- Validate dataset quality
Timeline: Weeks 1-2 (current)
Deliverable: BIOES dataset ready for training (train/val/test splits)
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.