ByteDance-Seed / ByteDance-Seed/SpatialTree
Bug: `RoboticArm` and `AgenticNavigation` sub-task labels are swapped (L4 / Goal-DrivenExecution)
- Lenguaje dominante
- Shell
- Estrellas
- 48
- Forks
- 0
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
Hi, thanks for open-sourcing this project — the benchmark design and task taxonomy are genuinely well organized, and the released resources are very helpful for the community. I really appreciate the effort your team put into building and maintaining this benchmark.
I noticed a possible labeling issue in the L4 / Goal-DrivenExecution tasks: it seems that the `RoboticArm` and `AgenticNavigation` sub-task labels may have been swapped. The current mappings appear inconsistent with the corresponding task contents and descriptions. It might be worth double-checking the annotations for these two categories.
| `spatree2` label | session_id | `metricfunc` | question text | actual task |
|---|---|---|---|---|
| `AgenticNavigation` | 5618–5867 (250) | `manipulateeval` | "…translation and rotation for the **robot arm's end-effector**… decompose the **end-effector movement** into 7 steps" | **robotic arm** |
| `RoboticArm` | 5868–6117 (250) | `agenticnaveval` | "**Task: Visual Navigation Action Sequence Generation** … expert **visual navigation agent** … **navigate a robot** …" | **navigation** |
Thanks again for releasing such a valuable benchmark and for your contributions to the community.
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Inspecciona las asignaciones de L4 / Goal-DrivenExecution para session_ids 5618–6117, incluidas las etiquetas de spatree2 y los valores de metricfunc mostrados en el informe. Compara cada asignación con el texto de su tarea, verifica si las dos etiquetas de subtarea están intercambiadas y actualiza las anotaciones si se confirma; después, vuelve a comprobar ambos rangos de 250 sesiones.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Área
- data, machine-learning
- Tipo de issue
- Error
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Estado de actividad
- Tranquilo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 58/100