alphacep / alphacep/vosk-api

Acoustic model training efficiency on noisy data

Offen
#1,366 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Jupyter Notebook
Sterne
15.1k
Forks
1.8k
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

Hello! I have a question.
There are several words (for example, 4 words and different forms of these words) that I want to catch in speech. Translation of other words is also needed, but its accuracy is not so significant. In normal audio, the translation of these words is good. But in audio with a lot of noise, fast speech, or a specific voice, these words quite often don't have a translation at all.

Would training an acoustic model on noisy data containing such words help in this case? Can the translation of these words be significantly improved? How big does the training audio set need to be for this?

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.