alphacep / alphacep/vosk-tts

Handling Mixed-Language Named Entities in Russian TTS Output

未关闭
#47 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Python
星标
274
派生
35
PR 合并指标
30 天内没有已合并 PR

描述

Hello,

I am researching Russian text-to-speech systems, particularly for cases where English named entities (like “KFC”, “Google”, etc.) are embedded in Russian sentences. Although I have not used the vosk-tts model directly yet, I have experience with other TTS models trained with VITS. In my experiments, those models could not generate natural-sounding audio for such mixed-language input—English words were pronounced using Russian phonetic rules, which affects intelligibility and naturalness.

I am interested in whether vosk-tts can handle these cases better, or if there are any plans to support more natural code-switching/mixed-language pronunciation in the future. Proper handling of named entities is very important for many real-world applications.

Do you have any suggestions or workarounds? Is this feature on the roadmap for vosk-tts? Any advice or pointers would be greatly appreciated.

Thank you for your work on this project!

Best regards,
afun

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。