alphacep / alphacep/vosk-api

Pass set of words to be recognized only does not work as expected

Ouverte
#1,335 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Jupyter Notebook
Étoiles
15.1k
Forks
1.8k
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

Hi
i would like to "concentrate" the recognition on certain words which are already part of the english/german models.
E.g the word "Whatsapp" is often recognized as "what's up". Its better for me to get nothing returned than "what's up".
I know that I can replace 'what's up' with 'Whatsapp' in postprocessing, but this defies the purpose.

I found in previous issues the following suggestions e.g. :

`KaldiRecognizer(model, 16000, "zero oh one two three four five six seven eight nine whatsapp")`

and

`KaldiRecognizer(model, wf.getframerate(), '[ "whatsapp", "[unk]" ]')`

So how i understand, this way you only get one of the words provided or nothing: https://github.com/alphacep/vosk-api/issues/107#issuecomment-756640282

This way i was hoping

A. to get a word e.g. "whatsapp" more "frequent" in my recognition result as it's specifically specified without the need to finetune the model
B. never the word "what's up" in my results as i didn't specify it in the set of words
C. better performance (?)
D. load the same model only once and only pass the words "to be concentrated on" to the `KaldiRecognizer ` function with every request

But in my case i get all possible words of the model recognized instead of the provided only.
It's pretty much the same as if i didn't provide any words at all. What am i doing wrong here?

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.