facebookresearch / facebookresearch/ProgramBench

csview: evaluator compatibility issues resolved — unchanged submission reaches official SOLVED (✅ 335 tests)

Aperta
#59 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
924
Fork
67
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

Update: we have now completed an end-to-end reproduction and isolated the evaluator-side causes affecting `wfxr__csview.8ac4de0`.

Using the same stabilized submission that previously produced an official score of 69, we identified and corrected two independent evaluator compatibility issues:

1. A test-environment dependency interaction caused one branch to fail before producing `results.xml`. The branch `run.sh` upgraded pytest at runtime, which produced an incompatibility with the installed `libtmux` pytest plugin. Preventing that unintended pytest upgrade restored normal branch execution.

2. JUnit testcase names produced under the evaluator's `eval.tests.*` namespace were compared literally against the corresponding `tests.*` names in `tests.json`. This caused valid executed tests to be simultaneously classified as unexpected and injected as `not_run`. We added namespace-aware matching while preserving the existing completeness checks.

We also retained the existing `pytest-timeout` compatibility handling (`thread` → `signal`) already required to prevent xdist worker loss on timed-out tests.

Verification sequence:

* isolated failing branch reproduced;
* minimal evaluator-side correction applied;
* isolated branch: 2/2 passed, no branch errors or warnings;
* evaluator regression suite: 41/41 passed;
* complete ProgramBench official evaluator rerun performed.

Final official result:

`wfxr__csview.8ac4de0 ✅ 335 tests`
`Average 100`

The resulting machine-readable evaluation artifact contains:

* `solution_branch: submission`
* `error_code: None`
* `error_details: None`
* 348 recorded test results
* 347 passed
* 1 skipped
* 0 `not_run`
* no branch errors
* no warnings

Most importantly, the submission archive was unchanged between the earlier score-69 run and the final solved run.

Submission SHA-256:

`a7b055b8dea5dddfadf88d381f634cdf2af3f6ac02b62f024740dd12b3bc512f`

Final official `eval.json` SHA-256:

`930daeeb3c4f080a98414b84338e2a6506f64e6e0ef3ab158d77be252a63d2d2`

ProgramBench repository HEAD used for the final run:

`963063c9271cc40fa179977356782ea4582e0b0c`

We preserved the complete evidence package: original submission, final `eval.json`, precheck reports, reconstruction state, evaluator state, evaluator diff, regression evidence, provenance, manifest, and SHA-256 inventory. The package was independently hash-verified after transfer from the Linux evaluation environment to Windows.

We would be happy to provide the evaluator patch and/or the verification package privately to the ProgramBench maintainers for independent reproduction.

Could you advise the preferred way to submit this result and evaluator finding for official independent verification?

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia esaminando il comportamento dell’evaluator descritto in relazione al branch run.sh, alla compatibilità pytest/libtmux, a tests.json e alla corrispondenza dei testcase con eval.tests.*. Esegui la suite di regressione dell’evaluator 41-test indicata e ispeziona l’eval.json finale; il lavoro è completato quando riproduci i risultati relativi alla compatibilità e verifichi in modo indipendente il risultato 335-test riportato, senza modificare la submission.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
testing-qa
Tipo di issue
Bug
Difficoltà
4/5
Tempo stimato
3-5 giorni
Stato di attività
Tranquilla
Chiarezza
Abbastanza chiara
Idoneità per principianti
35/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.