huggingface / huggingface/sentence-transformers

Unexpected high similarity

Open
#14 10 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
19.1k
Forks
2.9k
Avg merge
1d 19h
Merged PRs (30d)
61

Description

I am using bert-base-nli-stsb-mean-tokens model in an unsupervised fashion to get similarity between sentences.
It performs really good for some cases.
But on doing extensive analysis, I found some cases where such high score for similarity makes no sense.

I am trying to figure out why the similarity is so high for cases where sentences are extremely short or make no sense at all
What is really happening here?
Any leads would be helpful.

Thanks in advance,
for your reference
Screenshot 2019-08-22 at 6 37 33 PM

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the reported high-similarity results with the bert-base-nli-stsb-mean-tokens model, using the short or nonsensical sentence cases shown in the issue and its screenshot. Review the model's similarity behavior and document a supported explanation or a confirmed defect; no source file or test is named in the report.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.