allenai / allenai/OLMo

Question about Chinese Language Support and Model Retraining

未關閉
#875 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
type/question
主要語言
Python
星號
6.7k
分支
797
PR 合併指標
30 天內沒有已合併 PR

描述

### ❓ The question

I really appreciate the team's contribution in sharing this model for research and learning purposes. I have a question regarding its Chinese language capabilities. It appears the model is less optimized for Chinese inputs. Could you clarify:

1. What percentage of the pretraining corpus consists of Chinese data?
2. If I want to train a Chinese-optimized LLM based on this model, what technical recommendations would you suggest?
Thank you for your guidance.

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。