Unknown "words" in text.txt with Updating the language model
- Lingua principale
- Jupyter Notebook
- Stelle
- 15.1k
- Fork
- 1.8k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
I want to build a new grammar with a text.txt, all the commands are ok but the last one:
farcompilestrings --fst_type=compact --symbols=words.txt --keep_symbols text.txt | \
ngramcount | ngrammake | \
fstconvert --fst_type=ngram > Gr.new.fst
- If all the words in text.txt are in the words.txt => OK
- If there are "new words" in the text.txt (unknown words) => there are errors like:
`FATAL: FarCompileStrings: Compiling string number 2 in file text.txt failed with token_type = symbol and entry_type = line`
I read the -help and use the new command:
farcompilestrings --fst_type=compact --symbols=words.txt --unknown_symbol="" --keep_symbols text.txt | ngramcount | ngrammake | fstconvert --fst_type=ngram > Gr.new.fst
Another error raised:
`FATAL: FarCompileStrings: Label "-1" missing from symbol table: words.txt
FATAL: STListReader::STListReader: Wrong file type: `
I know that: "You can not introduce new words this way, that is something we will cover later.", but
Are there any ways to deal with "new words" in a big text?
Help me, plz!
Thanks in advance!
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.