aboutcode-org / aboutcode-org/scancode-toolkit

[GSoC] Train DeBERTa NER model for required phrase prediction

Đang mở
#5,137 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
GSoC
Ngôn ngữ chính
Python
Star
2.6k
Fork
791
Merge trung bình
1 ngày 12 giờ
Pull request đã merge (30 ngày)
5

Mô tả

continuation of #5077. uses the BIOES dataset produced there

## Goal
Finetune DeBERTa-v3-large to predict required phrase spans in license rule text using BIOES labels

## Tasks
- Subword tokenization with label alignment (-100 for continuation tokens)
- Training with class-weighted loss (handle O vs B/I/E/S imbalance)
- Include negative samples (rules without markers, all-O labels) for balanced training
- Evaluation: token F1, exact span match
- ONNX export for CPU inference

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.