ByteDance-Seed / ByteDance-Seed/SpatialTree
Bug: `RoboticArm` and `AgenticNavigation` sub-task labels are swapped (L4 / Goal-DrivenExecution)
- Lingua principale
- Shell
- Stelle
- 49
- Fork
- 0
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hi, thanks for open-sourcing this project — the benchmark design and task taxonomy are genuinely well organized, and the released resources are very helpful for the community. I really appreciate the effort your team put into building and maintaining this benchmark.
I noticed a possible labeling issue in the L4 / Goal-DrivenExecution tasks: it seems that the `RoboticArm` and `AgenticNavigation` sub-task labels may have been swapped. The current mappings appear inconsistent with the corresponding task contents and descriptions. It might be worth double-checking the annotations for these two categories.
| `spatree2` label | session_id | `metricfunc` | question text | actual task |
|---|---|---|---|---|
| `AgenticNavigation` | 5618–5867 (250) | `manipulateeval` | "…translation and rotation for the **robot arm's end-effector**… decompose the **end-effector movement** into 7 steps" | **robotic arm** |
| `RoboticArm` | 5868–6117 (250) | `agenticnaveval` | "**Task: Visual Navigation Action Sequence Generation** … expert **visual navigation agent** … **navigate a robot** …" | **navigation** |
Thanks again for releasing such a valuable benchmark and for your contributions to the community.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Direzione di ricerca
Esamina le mappature L4 / Goal-DrivenExecution per gli session_ids 5618–6117, incluse le etichette spatree2 e i valori di metricfunc mostrati nel report. Confronta ogni mappatura con il testo della relativa attività, verifica se le due etichette delle sotto-attività sono scambiate e aggiorna le annotazioni se ciò viene confermato; quindi ricontrolla entrambi gli intervalli di 250 sessioni.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Ambito
- data, machine-learning
- Tipo di issue
- Bug
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 58/100