huggingface / huggingface/setfit

for non default loss_class, onehot encoded labels dont work

Open
#403 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
2.8k
Forks
267
Avg merge
36m
Merged PRs (30d)
5

Description

One-hot encoded label for multi class dont work with non-default loss_class (eg:BatchHardTripletLoss, ...)
Note: Not using SetfitHead, but scikit learn one-vs-rest.

`TypeError: unhashable type: 'list'`

The error comes from line 353 in trainer.py `train_data_sampler = SentenceLabelDataset(train_examples, samples_per_label=self.samples_per_label)` which uses SentenceLabelDataset from sentence-transformer.

However when I change the labels to categorical enocoded it works.

This works:
dataset = Dataset.from_dict({"text": ["a 1", "b 1", "c 1", "a 2", "b 2"], "label": [0, 1, 2, 0, 1]})

This doesnt works:
dataset = Dataset.from_dict({"text": ["a 1", "b 1", "c 1", "a 2", "b 2", "c 2"], "label": [[1,0,0] [0,1,0] [0,0,1] [1,0,0] [0,1,0]]})

Maybe this is desired, but I couldn't find anywhere in documentation that this should be avoided, and example notebook talk about using One-hot encoded labels only.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.