allenai / allenai/dont-stop-pretraining

when do domain-adaptive pretraining, seems can not extend the vocabulary?

オープン
#36 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
544
フォーク
72
PR マージ指標
30日以内にマージされた PR はありません

説明

After use my own corpus to do domain-adaptive pretraining, the `vocab.txt` is the same size with the initialized model(BERT-base). In short, the domain-adaptive pretraining does not extend the vocabulary of the new domain? Therefore same specific
vocabulary of the new domain still not exist in the domain-adaptive pretraining result `vocab.txt`. Is that?

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。