bytedance / bytedance/SIFThinker

Question about depth reward

Open
#2 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
14
Forks
2
PR merge metrics
No merged PRs in 30d

Description

Hi, thank you for open-sourcing this great work. I have a quick clarification about the GRPO-SIF setup. In the current prompt, the model is asked to output interleaved and , but it does not seem to explicitly require a "depth" field inside each JSON item. Meanwhile, the depth_consistency reward appears to depend on parsing that depth value. Could this mismatch be the reason why depth reward is often 0? I would really appreciate your guidance on whether this is expected or if the prompt should explicitly enforce depth output.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.