Question about Chinese Language Support and Model Retraining
Open
type/question
- Dominant language
- Python
- Stars
- 6.7k
- Forks
- 797
- PR merge metrics
- No merged PRs in 30d
Description
### ❓ The question
I really appreciate the team's contribution in sharing this model for research and learning purposes. I have a question regarding its Chinese language capabilities. It appears the model is less optimized for Chinese inputs. Could you clarify:
1. What percentage of the pretraining corpus consists of Chinese data?
2. If I want to train a Chinese-optimized LLM based on this model, what technical recommendations would you suggest?
Thank you for your guidance.
Contributor guide
Assessment
This issue has not been assessed yet.