allenai / allenai/OLMo

Question about Chinese Language Support and Model Retraining

Open
#875 2 comments 0 reactions 0 assignees View on GitHub
type/question
Dominant language
Python
Stars
6.7k
Forks
797
PR merge metrics
No merged PRs in 30d

Description

### ❓ The question

I really appreciate the team's contribution in sharing this model for research and learning purposes. I have a question regarding its Chinese language capabilities. It appears the model is less optimized for Chinese inputs. Could you clarify:

1. What percentage of the pretraining corpus consists of Chinese data?
2. If I want to train a Chinese-optimized LLM based on this model, what technical recommendations would you suggest?
Thank you for your guidance.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.