modelscope / modelscope/ms-swift
qwen3-vl grounding 任务,GRPO数据集应该如何准备?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 15.7k
- Forks
- 1.7k
- Avg merge
- 1d 16h
- Merged PRs (30d)
- 136
Description
Checklist / 检查清单
- I have searched existing issues, and this is a new question or discussion topic. / 我已经搜索过现有的 issues,确认这是一个新的问题与讨论。
Question Description / 问题描述
我做了两版:
1、bbox直接当做普通文本。数据集样例如下(普通方式预测处理训练是没问题的,但是不是缺失了bbox的特殊字符):
{ "messages": [ { "content": "# 角色:\n你是一位专业的零售电器门店陈列核查员。你的任务是对比“标准陈列图”与“待检查图”,识别待检查图中出现的空柜/空展区现象。\n\n# 核心原则:\n1. **比对粒度**:你的判断唯一依据是“标准陈列图”与“待检查图”的像素级差异。\n2. **排除干扰**:忽略人员变化、光照亮度、电子屏幕、宣传物料的等差异带来的噪音,仅聚焦于“商品实体”是否缺失并造成了显著的“空柜/空展区”现象。\n\n# 输出要求:\n- **空值处理**:若全图无空柜/空展区现象,直接输出 []。\n- **格式要求**:必须严格遵守 JSON 格式,参考如下:\n[\n{\"bbox\": [x, x, x, x], \"label\": \"blank\"},\n{\"bbox\": [x, x, x, x], \"label\": \"blank\"},...\n]\n\n# 待检查任务:\n现在,请开始仔细对比当前上传的两张图片。\n**标准陈列图**:<image>\n**待检查图**:<image>。", "role": "user" }, { "content": "<reason>\n空,此处无需think过程</reason>\n<answer>\n[{\"bbox\": [760, 448, 996, 996], \"label\": \"blank\"}, {\"bbox\": [2, 637, 236, 995], \"label\": \"blank\"}]</answer>", "role": "assistant" } ], "images": [ "/llm/wjw/LLM_train/LLaMA-Factory/data_utils/空柜检测/train_images_0122/original_images/2d5f78f9-ea28-4f73-b0bd-004e22713a5d_A.jpg", "/llm/wjw/LLM_train/LLaMA-Factory/data_utils/空柜检测/train_images_0122/processed_images/2d5f78f9-ea28-4f73-b0bd-004e22713a5d_B.jpg" ], "solution": "<reason>\n空,此处无需think过程</reason>\n<answer>\n[{\"bbox\": [760, 448, 996, 996], \"label\": \"blank\"}, {\"bbox\": [2, 637, 236, 995], \"label\": \"blank\"}]</answer>" },
2、参考官方说明文档的模式处理(这种在设计ORM时,solution部分,传到ORM中的的依然是普通"",无法获得具体的bbox计算ORM):
{ "messages": [ { "role": "user", "content": "# 角色:\n你是一位专业的零售电器门店陈列核查员。你的任务是对比“标准陈列图”与“待检查图”,识别待检查图中出现的空柜/空展区现象。\n\n# 核心原则:\n1. **比对粒度**:你的判断唯一依据是“标准陈列图”与“待检查图”的像素级差异。\n2. **排除干扰**:忽略人员变化、光照亮度、电子屏幕、宣传物料的等差异带来的噪音,仅聚焦于“商品实体”是否缺失并造成了显著的“空柜/空展区”现象。\n\n# 输出要求:\n- 请直接输出所有检测到的“空柜/空展区”的检测框(Bounding Boxes)。\n- 若全图无空柜/空展区现象,请输出:[]。\n\n# 待检查任务:\n现在,请开始仔细对比当前上传的两张图片。\n**标准陈列图**:<image>\n**待检查图**:<image>。" }, { "role": "assistant", "content": "<bbox><bbox>" } ], "task": "blank_inspection", "solution": "<bbox><bbox>", "images": [ "/llm/wjw/LLM_train/LLaMA-Factory/data_utils/空柜检测/train_images_0122/original_images/2d5f78f9-ea28-4f73-b0bd-004e22713a5d_A.jpg", "/llm/wjw/LLM_train/LLaMA-Factory/data_utils/空柜检测/train_images_0122/processed_images/2d5f78f9-ea28-4f73-b0bd-004e22713a5d_B.jpg" ], "objects": { "ref": [ "blank", "blank" ], "bbox": [ [ 0.76, 0.448, 0.996, 0.996 ], [ 0.002, 0.637, 0.236, 0.995 ] ], "bbox_type": "norm1", "image_id": [ 1, 1 ] } },
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Compare the two dataset examples with the official grounding-data documentation, then trace how the ORM receives the assistant content and solution fields. Confirm whether the expected normalized bounding boxes and labels are preserved rather than reduced to literal markers; done means a documented format that lets GRPO training and ORM processing recover the boxes correctly.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100