alphacep / alphacep/vosk-api

German small model and the umlaut

Open
#2,040 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
15.1k
Forks
1.8k
PR merge metrics
No merged PRs in 30d

Description

Hello, folks.

I am trying out the small German model "vosk-model-small-de-0.15" by passing in grammar to the recognizer. What I am experiencing is that it works great, except for when a word that is passed in grammar contains an umlaut.

For instance, passing in the following json:
"[\"eins\",\"zwei\",\"drei\",\"vier\",\"fünf\",\"sechs\",\"sieben\",\"acht\",\"neun\",\"zehn\",\"elf\",\"zwölf\"]"

The recognizer will respond perfectly with everything but "fünf" and "zwölf".

I tried with the large model "vosk-model-de-0.21", and it seems to work with words that contain an umlaut. Is there something I am missing in regard to the small model or is this a known limitation? I haven't been able to find information on this, so I am hoping it is just me ;)

Thank you for all of your hard work on this.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.