alphacep / alphacep/vosk-api

Pass set of words to be recognized only does not work as expected

オープン
#1,335 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Jupyter Notebook
スター
15.1k
フォーク
1.8k
PR マージ指標
30日以内にマージされた PR はありません

説明

Hi
i would like to "concentrate" the recognition on certain words which are already part of the english/german models.
E.g the word "Whatsapp" is often recognized as "what's up". Its better for me to get nothing returned than "what's up".
I know that I can replace 'what's up' with 'Whatsapp' in postprocessing, but this defies the purpose.

I found in previous issues the following suggestions e.g. :

`KaldiRecognizer(model, 16000, "zero oh one two three four five six seven eight nine whatsapp")`

and

`KaldiRecognizer(model, wf.getframerate(), '[ "whatsapp", "[unk]" ]')`

So how i understand, this way you only get one of the words provided or nothing: https://github.com/alphacep/vosk-api/issues/107#issuecomment-756640282

This way i was hoping

A. to get a word e.g. "whatsapp" more "frequent" in my recognition result as it's specifically specified without the need to finetune the model
B. never the word "what's up" in my results as i didn't specify it in the set of words
C. better performance (?)
D. load the same model only once and only pass the words "to be concentrated on" to the `KaldiRecognizer ` function with every request

But in my case i get all possible words of the model recognized instead of the provided only.
It's pretty much the same as if i didn't provide any words at all. What am i doing wrong here?

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。