allenai / allenai/specter

Using trained model: which tokenizer?

Open
#38 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
592
Forks
58
PR merge metrics
No merged PRs in 30d

Description

I have trained a new model following the guildeline in `README.md`. The model was trained on **my own dataset** of scientific articles. Now, in order to use the trained model, I need a tokenizer. Which one should I use? Do I need to load the vocabulary from disk in case the vocabulary used during training is different from pretrained ones?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.