aboutcode-org / aboutcode-org/scancode-toolkit

[GSoC] Build NER training dataset from annotated rules

Aperta
#5,077 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
GSoC
Lingua principale
Python
Stelle
2.6k
Fork
791
Merge medio
1g 12h
PR unite (30g)
5

Descrizione

Tasks:
- Parse .RULE files and extract extract required phrase markers
- Propagate annotations (with CLI commands from #3924 ) to unannotated rules (with matching text) to expand the dataset
- Improve composite rules (AND/OR expressions where some aren't marked)
- Convert to BIOES labels
- Split into train/val/test by rule family
- Validate dataset quality

Timeline: Weeks 1-2 (current)

Deliverable: BIOES dataset ready for training (train/val/test splits)

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.