THUDM / THUDM/slime

SFT loss mask

Open
#514 0 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
8.5k
Forks
1.3k
Avg merge
5h 36m
Merged PRs (30d)
22

Description

I'm using slime for GLM-4.5-AIR SFT training.
Since GLM-4.5 doesn’t use a dedicated end-of-turn token, the next <|user|> or <|observation|> tag should serve as the end marker for the current assistant response.

I noticed that in the function gen_multi_turn_loss_mask_qwen, the entire user segment’s loss_mask is set to 0.
Would this cause the model to never learn the loss associated with generating the end-of-turn boundary (i.e., the transition to the next <|user|> or <|observation|>) during training?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by inspecting the gen_multi_turn_loss_mask_qwen function and the SFT loss-mask handling for GLM-4.5-AIR. Trace how the next <|user|> or <|observation|> boundary is represented, then determine whether the current mask includes the assistant-to-boundary transition. Done means the expected boundary-loss behavior is established and the issue has a clearly defined follow-up change or explanation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.