kaldi-asr / kaldi-asr/kaldi

Help with training own language model

Open
#4,835 0 comments 0 reactions 0 assignees View on GitHub
discussion
Dominant language
Shell
Stars
15.5k
Forks
5.4k
PR merge metrics
No merged PRs in 30d

Description

I am interested in training my own Bangla language model, but I'm not sure where to start. I have my own Bangla dataset with audio and text, and I would like to use it to train a model that can transcribe Bangla speech to text offline. I am looking for guidance on how to preprocess my data, train a model, and evaluate its performance.

Can someone please provide detailed instructions or point me to a tutorial or guide that can help me with this process? Here are some specific questions I have:

- What are the best practices for preprocessing Bangla audio and text data?
- How do I create a Bangla language model and generate the necessary files for training a model?
- What are the recommended training parameters and settings for training a Bangla language model ?
- How do I evaluate the performance of my trained model, and what metrics should I use?
Installation guidelines from scratch.

I would appreciate any help or advice that can be provided. Thank you in advance!

Contributor guide

No contributing guide indexed for this repository

Research direction

No files or tests are identified. Start with Kaldi's installation and speech-recognition training documentation, then determine how the Bangla audio and text dataset should be prepared, how training should be configured, and which evaluation metrics and offline transcription results would demonstrate completion.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.