SkyworkAI / SkyworkAI/UniPic

Training data composition

Open
#25 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
875
Forks
43
PR merge metrics
No merged PRs in 30d

Description

Nice work. I have another question Are the 130M pairs of data used in the three stages of PT Alignment CT the same? What is the ratio of understanding to generation data?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue asks about the 130M training pairs across the PT, Alignment, and CT stages, but names no files, tests, or entry points. Start by locating the project’s training-data definitions or documentation for those stages and compare their understanding-to-generation composition. Done means documenting whether the datasets are shared and stating the ratio for each relevant stage.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.