Question about Chinese Language Support and Model Retraining
未關閉
type/question
- 主要語言
- Python
- 星號
- 6.7k
- 分支
- 797
- PR 合併指標
- 30 天內沒有已合併 PR
描述
### ❓ The question
I really appreciate the team's contribution in sharing this model for research and learning purposes. I have a question regarding its Chinese language capabilities. It appears the model is less optimized for Chinese inputs. Could you clarify:
1. What percentage of the pretraining corpus consists of Chinese data?
2. If I want to train a Chinese-optimized LLM based on this model, what technical recommendations would you suggest?
Thank you for your guidance.
貢獻指南
評估
這個 Issue 還沒有評估資料。