modelscope / modelscope/ms-swift

diffusion gemma 微调时数据输出tokens的数量可否支持 > config 中的 256 tokens?

Open
#9,739 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
15.7k
Forks
1.7k
Avg merge
1d 16h
Merged PRs (30d)
136

Description

Checklist / 检查清单
  • I have searched existing issues, and this is a new feature request. / 我已经搜索过现有的 issues,确认这是一个新的 Feature Request。
Feature Request Description / Feature Request 描述

目前正如 diffusion_gemma.sh 脚本所注释的文本,the response length of a single sample must be less than config.canvas_length,数据中模型回复的长度不超过 canvas_length 也就是256tokens,但是很多时候训练集超过了这个长度,是否可以支持超过这个长度的模型训练呢?

Pull Request / Pull Request 信息

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the comments and configuration used by diffusion_gemma.sh, especially canvas_length, and trace where the 256-token response limit is enforced during fine-tuning. Determine which related data-processing and training paths depend on that limit. Done means samples with responses longer than 256 tokens can be trained without truncation or failure, with the existing script flow still working.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.