jina-ai / jina-ai/node-DeepResearch

Questions for Query Deduplication

Open
#112 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
TypeScript
Stars
5.2k
Forks
461
PR merge metrics
No merged PRs in 30d

Description

Hello, nice work on this DeepResearch project.

In the [blog](https://jina.ai/news/a-practical-guide-to-implementing-deepsearch-deepresearch/), it says that -
```
For query deduplication, we initially used an LLM-based solution, but found it difficult to control the similarity threshold. We eventually switched to [jina-embeddings-v3](https://jina.ai/?sui&model=jina-embeddings-v3), which excels at semantic textual similarity tasks. This enables cross-lingual deduplication without worrying that non-English queries would be filtered. The embedding model ended up being crucial not for memory retrieval as initially expected, but for efficient deduplication.
```
However, in the code, it still seems to use LLM for deduplication.
https://github.com/jina-ai/node-DeepResearch/blob/main/src/tools/dedup.ts#L44

Were there some trade-offs you observed?

Thank you.

- Youngjoon Jang

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.