allenai / allenai/vla-evaluation-harness

BEHAVIOR Challenge 2026: integrate the official evaluator through a protocol bridge

Ouverte
#114 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Python
Étoiles
606
Forks
49
Merge moyen
3 j 11 h
PR mergées (30 j)
7

Description

Follow-up to #113, whose immediate request (label the current integration as BEHAVIOR 2025) is handled in the linked PR. This issue tracks the actual 2026 evaluation path.

## Problem

The official 2026 BEHAVIOR Challenge pins BEHAVIOR-1K v3.9.1, evaluates 100 tasks, and drives the policy server through `python -m omnigibson.eval.eval`: the evaluator waits for `/healthz`, sends flattened observations over WebSocket, and expects a msgpack response containing an `action` array. Official RGB + depth evaluation uses `omnigibson.eval.wrappers.RGBDFullResWrapper`, and reported results use instance indices 0-9. Instructions: https://behavior.stanford.edu/challenge/evaluation.html

Porting the harness's `Behavior1KBenchmark` to v3.9.1 would not establish 2026 submission compatibility, because the harness would still re-implement the evaluator and its own episode loop.

## Proposed approach: protocol bridge

Keep the official evaluator as the driver and bridge its wire protocol to the harness model-server protocol:

- New `docker/Dockerfile.behavior1k2026` pinned to BEHAVIOR-1K v3.9.1 and the Isaac Sim version it requires, separate from the existing 2025 image.
- A thin shim server that exposes the official policy-server interface (`/healthz`, flattened observations in, msgpack `{action: [...]}` out) and translates each request to the vla-eval model-server WebSocket protocol.
- The benchmark entry runs `python -m omnigibson.eval.eval` (100 tasks, instances 0-9) inside the container and collects results from the evaluator's output.

This validates the real submission path end to end and reuses every existing model server unchanged. A passing run is then a true 2026 compatibility claim, which the 2025 adapter cannot provide.

## Scope notes

- The 2025 integration stays as is, now explicitly labeled.
- #54 (BEHAVIOR-1K Docker image support) overlaps with the image work here.

Guide de contribution

Ouvrir le guide de contribution

Piste de recherche

Start from the proposed new docker/Dockerfile.behavior1k2026 and the official entry point python -m omnigibson.eval.eval. Read the BEHAVIOR 2026 evaluator protocol: /healthz, flattened observations over WebSocket, and msgpack {action: [...]} responses. Done means the official evaluator can run the 100 tasks for instances 0-9 through the bridge and collect its results output.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
docker, python
Domaine
ai, backend, devops
Type d'issue
Fonctionnalité
Difficulté
5/5
Temps estimé
Plus d'une semaine
Activité
Calme
Clarté
Plutôt claire
Accessibilité débutants
28/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.