facebookresearch / facebookresearch/ProgramBench
ProgramBench csview: false not_run records and missing results.xml
- Lingua principale
- Python
- Stelle
- 928
- Fork
- 66
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
### Environment
- ProgramBench: `1.2.4`
- Repository commit: `963063c9271cc40fa179977356782ea4582e0b0c`
- Instance: `wfxr__csview.8ac4de0`
- Cleanroom image: `programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6`
- Image digest: `sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd`
- Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
- Inference container: `--network none`
### Reproduction
The same submission was evaluated twice:
```bash
uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5
uv run programbench eval \
~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
--docker-cpus 1 \
--branch-retries 3
```
Both runs produced the same raw result counts:
- passed: 344
- failed: 1
- skipped: 1
- not_run: 155
- displayed score: 68
The generated executable SHA-256 was identical in both evaluations:
`a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888`
### Issue 1: test-name prefix mismatch creates 153 false `not_run` records
For branch `d31025219260`, the evaluator reports both:
- `153/153 expected tests missing from JUnit XML`
- `153 test(s) in JUnit XML not in tests.json`
The records form 153 exact pairs after removing only the leading `eval.` prefix. For example:
- metadata: `tests.test_edge_cases.test_all_cells_empty`
- JUnit: `eval.tests.test_edge_cases.test_all_cells_empty`
All 153 JUnit-side records passed, while the metadata-side names were counted as `not_run`.
Suggested fix: normalize an optional leading `eval.` prefix before matching JUnit names to `tests.json`, or generate both sources with the same package root.
### Issue 2: branch `efa8c407dbe3` never produces `results.xml`
The branch fails with:
```text
results_read_failed: Could not find the file /workspace/eval/results.xml
```
This occurred in the default run and again in the serial run. With `--branch-retries 3`, the initial attempt plus all three retries failed identically. Two expected tests remain `not_run` because no JUnit result file is available.
Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when `results.xml` is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.
### Issue 3: `--help` golden differs from the cleanroom reference behavior
The only completed-test failure is:
`eval.tests.test_cli_basics.test_help_exact`
However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:
```bash
diff -u \
<(/workspace/executable --help) \
<(/candidate/csview_reconstruction_v0_5/executable --help)
```
The local cleanroom differential suite also passed `49/49`, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in `task_cleanroom_v6`.
Suggested fix: regenerate `help.txt` directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.
### Why this matters
Among completed scored tests, the observed pass rate is `344 / 345 = 99.710145%`. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.
This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.
### Public evidence attached
The attached public-safe evidence archive contains:
- both raw evaluation JSON files;
- a machine-readable audit summary;
- the deterministic audit script;
- the full bilingual technical report;
- SHA-256 checksums.
No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.
[programbench_csview_public_issue_evidence_v1.zip](https://github.com/user-attachments/files/30601982/programbench_csview_public_issue_evidence_v1.zip)
Guida per i contributori
Apri la guida per i contributori
Direzione di ricerca
Inizia ispezionando il JSON di valutazione grezzo allegato, lo script di audit e il rapporto tecnico, quindi confronta tests.json con i nomi JUnit e esamina il caso di results.xml mancante. Verifica il riferimento in help.txt rispetto all’eseguibile conservato. Il lavoro è completato quando i test corrispondenti vengono conteggiati una sola volta, i fallimenti espongono l’output del runner e producono risultati utilizzabili, e il golden di help è associato al riferimento dichiarato.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- docker, python
- Ambito
- testing, tooling
- Tipo di issue
- Bug
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Stato di attività
- Tranquilla
- Chiarezza
- Abbastanza chiara
- Idoneità per principianti
- 48/100