语义检索bad case,是模型的问题吗,需要怎么解决
Open
- Dominant language
- C++
- Stars
- 2.6k
- Forks
- 657
- PR merge metrics
- No merged PRs in 30d
Description
用AnyQ自带的模型检索小猪佩奇,命中的相似文本
1. 周星驰新版女喜剧之王!
2. 汉堡排骨米饭
3. 耶稣我爱你~C调
4. 小猪佩奇全集 煎饼
5. 心连心爱你一生,蝴蝶兰之紫水晶
6. 熊出没 熊出没之秋人团团转光头强训练营2
7. 黄梅戏“一唱两走”天地宽
8. 小猪佩奇全集 牙仙子
9. 凉拌豆皮超辣
10. 江小M 直到黎明 萧D哥上场!
词典里面有小猪佩奇,但是analysis的时候就切成了小猪和佩奇,可能是因为这个,样本里面明明很多小猪佩奇相关的,但是召回了很多一点都不相关的,是索引的问题还是模型的问题?怎么解决呢?
Contributor guide
No contributing guide indexed for this repository
Research direction
No source file or test is named. Reproduce the “小猪佩奇” query and inspect the analysis tokenization, dictionary matching, index retrieval, and model-ranking outputs to determine which stage causes the unrelated results; done means identifying the responsible stage and documenting a validated corrective approach.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning, search
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100