Request to Modify Code to Enable TEXT_SPLITTER_EMBEDDING_MODEL Customization through Configuration File
@sumitkbh 已经在做这个了。
开始于 2024年1月18日。
评估
这个 Issue 还没有评估数据。
描述
I am looking to create a Chinese RAG demo service using RetrievalAugmentedGeneration.
However, I encountered an issue where the default SentenceTransformersTokenTextSplitter model used in the RetrievalAugmentedGeneration/common/utils.py file is hardcoded as 'intfloat/e5-large-v2'. This model generates a significant number of [UNK] tokens when processing Chinese text.
I would like the ability to specify a specific model for the text splitter, similar to how the embedding model can be specified through the config.yaml file.
Thank you for your assistance and support.
- 主要语言
- Jupyter Notebook
- 星标
- 4.2k
- 派生
- 1.1k
- 平均合并
- 10 小时 15 分钟
- 30 天内合并 PR
- 1
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/GenerativeAIExamples 的其他 Issue
-
难度 1/5 1 小时以内 新手友好度 70/100
NVIDIA/GenerativeAIExamples#361 · 2 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 72/100
NVIDIA/GenerativeAIExamples#299 ·
-
难度 3/5 1-2 天 新手友好度 35/100
NVIDIA/GenerativeAIExamples#437 ·
-
难度 1/5 1 小时以内 新手友好度 45/100
NVIDIA/GenerativeAIExamples#400 ·
-
难度 5/5 一周以上 新手友好度 10/100
NVIDIA/GenerativeAIExamples#399 ·