alphacep / alphacep/vosk-api

Unknown "words" in text.txt with Updating the language model

Abierto
#612 5 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Jupyter Notebook
Estrellas
15.1k
Forks
1.8k
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

I want to build a new grammar with a text.txt, all the commands are ok but the last one:
farcompilestrings --fst_type=compact --symbols=words.txt --keep_symbols text.txt | \
ngramcount | ngrammake | \
fstconvert --fst_type=ngram > Gr.new.fst
- If all the words in text.txt are in the words.txt => OK
- If there are "new words" in the text.txt (unknown words) => there are errors like:
`FATAL: FarCompileStrings: Compiling string number 2 in file text.txt failed with token_type = symbol and entry_type = line`
I read the -help and use the new command:
farcompilestrings --fst_type=compact --symbols=words.txt --unknown_symbol="" --keep_symbols text.txt | ngramcount | ngrammake | fstconvert --fst_type=ngram > Gr.new.fst
Another error raised:
`FATAL: FarCompileStrings: Label "-1" missing from symbol table: words.txt
FATAL: STListReader::STListReader: Wrong file type: `
I know that: "You can not introduce new words this way, that is something we will cover later.", but
Are there any ways to deal with "new words" in a big text?
Help me, plz!
Thanks in advance!

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.