Training data composition
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 875
- Forks
- 43
- PR merge metrics
- No merged PRs in 30d
Description
Nice work. I have another question Are the 130M pairs of data used in the three stages of PT Alignment CT the same? What is the ratio of understanding to generation data?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue asks about the 130M training pairs across the PT, Alignment, and CT stages, but names no files, tests, or entry points. Start by locating the project’s training-data definitions or documentation for those stages and compare their understanding-to-generation composition. Done means documenting whether the datasets are shared and stating the ratio for each relevant stage.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100