baidu / baidu/lac

增量训练是否会使得原分词/标注效果变差

Open
#70 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

例如使用1个垂域的少量数据,大概几万条样本,进行增量训练。
是否会出现原分词/标注效果变差、而仅retrain的垂域效果好的问题?

Contributor guide

No contributing guide indexed for this repository

Research direction

No file, test, or entry point is named. Start by locating the incremental-training workflow and compare the original model's segmentation and tagging results with the domain-specific results. Done means producing a documented, reproducible answer about whether the original capabilities degrade.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.