ByteDance-Seed / ByteDance-Seed/SpatialTree

Bug: `RoboticArm` and `AgenticNavigation` sub-task labels are swapped (L4 / Goal-DrivenExecution)

Offen
#1 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Shell
Sterne
49
Forks
0
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

Hi, thanks for open-sourcing this project — the benchmark design and task taxonomy are genuinely well organized, and the released resources are very helpful for the community. I really appreciate the effort your team put into building and maintaining this benchmark.

I noticed a possible labeling issue in the L4 / Goal-DrivenExecution tasks: it seems that the `RoboticArm` and `AgenticNavigation` sub-task labels may have been swapped. The current mappings appear inconsistent with the corresponding task contents and descriptions. It might be worth double-checking the annotations for these two categories.

| `spatree2` label | session_id | `metricfunc` | question text | actual task |
|---|---|---|---|---|
| `AgenticNavigation` | 5618–5867 (250) | `manipulateeval` | "…translation and rotation for the **robot arm's end-effector**… decompose the **end-effector movement** into 7 steps" | **robotic arm** |
| `RoboticArm` | 5868–6117 (250) | `agenticnaveval` | "**Task: Visual Navigation Action Sequence Generation** … expert **visual navigation agent** … **navigate a robot** …" | **navigation** |

Thanks again for releasing such a valuable benchmark and for your contributions to the community.

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Untersuche die L4 / Goal-DrivenExecution-Zuordnungen für session_ids 5618–6117, einschließlich der spatree2-Labels und der im Bericht angezeigten metricfunc-Werte. Vergleiche jede Zuordnung mit ihrem Aufgabentext, überprüfe, ob die beiden Subtask-Labels vertauscht sind, und aktualisiere die Annotationen, falls dies bestätigt wird; überprüfe anschließend erneut beide Bereiche mit jeweils 250 Sessions.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Bereich
data, machine-learning
Issue-Typ
Bug
Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Aktivitätsstatus
Ruhig
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
58/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.