huggingface / huggingface/sentence-transformers
Unexpected high similarity
- Dominant language
- Python
- Stars
- 19.1k
- Forks
- 2.9k
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 61
Description
I am using bert-base-nli-stsb-mean-tokens model in an unsupervised fashion to get similarity between sentences.
It performs really good for some cases.
But on doing extensive analysis, I found some cases where such high score for similarity makes no sense.
I am trying to figure out why the similarity is so high for cases where sentences are extremely short or make no sense at all
What is really happening here?
Any leads would be helpful.
Thanks in advance,
for your reference
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reproducing the reported high-similarity results with the bert-base-nli-stsb-mean-tokens model, using the short or nonsensical sentence cases shown in the issue and its screenshot. Review the model's similarity behavior and document a supported explanation or a confirmed defect; no source file or test is named in the report.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100