huggingface / huggingface/hmtl

Issue with only train data being used for vocab creation

Open
#15 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.2k
Forks
144
PR merge metrics
No merged PRs in 30d

Description

Hi team,
Thanks for this wonderful repo . The code in the repo is generic and can easily be reused. I wanted to ask that during creation of the vocab in all the models only training tokens are being used.
```
"datasets_for_vocab_creation": ["train"]
```
So in cases when we are using the multitask model we have a large coverage of tokens as we have a large vocab that consists of tokens from all datasets. So there is a high probability of test token to be found in that vocab. Whereas in case of using only single model the vocab size is less and there is a large chance of a token being OOV (Out of Vocab).
So how do we make sure that the improvements are due to multitask learning rather then due to large coverage of vocabulary in case of multitask learning?

The other point was that if we only consider vocab made from training data we make our model work well on only tokens that are present in training data which makes us loose important token information that is present in the word embeddings for those tokens which are not present in the training data.

It would be great to hear your thoughts on it.

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names the datasets_for_vocab_creation setting but no source file or test; start by locating that setting and the vocabulary-creation path for single-task and multitask models. Compare how train-only vocabularies and broader dataset vocabularies are configured, then document the maintainer-approved evaluation approach and verify it in the relevant model runs.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.