alphacep / alphacep/vosk-api

Unknown "words" in text.txt with Updating the language model

オープン
#612 コメント 5 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Jupyter Notebook
スター
15.1k
フォーク
1.8k
PR マージ指標
30日以内にマージされた PR はありません

説明

I want to build a new grammar with a text.txt, all the commands are ok but the last one:
farcompilestrings --fst_type=compact --symbols=words.txt --keep_symbols text.txt | \
ngramcount | ngrammake | \
fstconvert --fst_type=ngram > Gr.new.fst
- If all the words in text.txt are in the words.txt => OK
- If there are "new words" in the text.txt (unknown words) => there are errors like:
`FATAL: FarCompileStrings: Compiling string number 2 in file text.txt failed with token_type = symbol and entry_type = line`
I read the -help and use the new command:
farcompilestrings --fst_type=compact --symbols=words.txt --unknown_symbol="" --keep_symbols text.txt | ngramcount | ngrammake | fstconvert --fst_type=ngram > Gr.new.fst
Another error raised:
`FATAL: FarCompileStrings: Label "-1" missing from symbol table: words.txt
FATAL: STListReader::STListReader: Wrong file type: `
I know that: "You can not introduce new words this way, that is something we will cover later.", but
Are there any ways to deal with "new words" in a big text?
Help me, plz!
Thanks in advance!

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。