aboutcode-org / aboutcode-org/scancode-toolkit

[GSoC] Train DeBERTa NER model for required phrase prediction

Aperta
#5,137 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
GSoC
Lingua principale
Python
Stelle
2.6k
Fork
791
Merge medio
1g 12h
PR unite (30g)
5

Descrizione

continuation of #5077. uses the BIOES dataset produced there

## Goal
Finetune DeBERTa-v3-large to predict required phrase spans in license rule text using BIOES labels

## Tasks
- Subword tokenization with label alignment (-100 for continuation tokens)
- Training with class-weighted loss (handle O vs B/I/E/S imbalance)
- Include negative samples (rules without markers, all-O labels) for balanced training
- Evaluation: token F1, exact span match
- ONNX export for CPU inference

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.