facebookresearch / facebookresearch/ProgramBench

ProgramBench csview: false not_run records and missing results.xml

Open
#56 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
924
Forks
67
PR merge metrics
No merged PRs in 30d

Description

### Environment

- ProgramBench: `1.2.4`
- Repository commit: `963063c9271cc40fa179977356782ea4582e0b0c`
- Instance: `wfxr__csview.8ac4de0`
- Cleanroom image: `programbench/wfxr_1776_csview.8ac4de0:task_cleanroom_v6`
- Image digest: `sha256:e885e345ca055180d8d5579df46cc62eedf818f00aeab5712f4a6f456eb92fbd`
- Host: Ubuntu 24.04.4 LTS, x86_64, native Docker
- Inference container: `--network none`

### Reproduction

The same submission was evaluated twice:

```bash
uv run programbench eval ~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5

uv run programbench eval \
~/SINTEZ/PROGRAMBENCH_EVALS/csview_v0_5_serial \
--docker-cpus 1 \
--branch-retries 3
```

Both runs produced the same raw result counts:

- passed: 344
- failed: 1
- skipped: 1
- not_run: 155
- displayed score: 68

The generated executable SHA-256 was identical in both evaluations:

`a86228c9fd2d3c8d8dc25fa409418b8e48115b9dbf0ed5072b8797afa1e3d888`

### Issue 1: test-name prefix mismatch creates 153 false `not_run` records

For branch `d31025219260`, the evaluator reports both:

- `153/153 expected tests missing from JUnit XML`
- `153 test(s) in JUnit XML not in tests.json`

The records form 153 exact pairs after removing only the leading `eval.` prefix. For example:

- metadata: `tests.test_edge_cases.test_all_cells_empty`
- JUnit: `eval.tests.test_edge_cases.test_all_cells_empty`

All 153 JUnit-side records passed, while the metadata-side names were counted as `not_run`.

Suggested fix: normalize an optional leading `eval.` prefix before matching JUnit names to `tests.json`, or generate both sources with the same package root.

### Issue 2: branch `efa8c407dbe3` never produces `results.xml`

The branch fails with:

```text
results_read_failed: Could not find the file /workspace/eval/results.xml
```

This occurred in the default run and again in the serial run. With `--branch-retries 3`, the initial attempt plus all three retries failed identically. Two expected tests remain `not_run` because no JUnit result file is available.

Suggested fix: preserve and expose the branch container's test-runner stdout/stderr when `results.xml` is missing, and ensure the runner emits a minimal JUnit file even on setup or collection failure.

### Issue 3: `--help` golden differs from the cleanroom reference behavior

The only completed-test failure is:

`eval.tests.test_cli_basics.test_help_exact`

However, in the official cleanroom container, a byte-for-byte comparison between the preserved reference executable and the reconstructed executable produced no diff:

```bash
diff -u \
<(/workspace/executable --help) \
<(/candidate/csview_reconstruction_v0_5/executable --help)
```

The local cleanroom differential suite also passed `49/49`, including the exact help behavior. This suggests that the branch golden may not correspond to the preserved reference executable in `task_cleanroom_v6`.

Suggested fix: regenerate `help.txt` directly from the preserved reference executable used for this instance, or document which reference/version the golden represents.

### Why this matters

Among completed scored tests, the observed pass rate is `344 / 345 = 99.710145%`. The displayed score of 68 is materially affected by evaluator/metadata artifacts: 153 prefix-mismatched names and two tests in a branch that never emits results.

This is not a request to rewrite an official score. It is a request to inspect and correct the evaluation artifacts so that the benchmark result reflects the tests that actually ran.

### Public evidence attached

The attached public-safe evidence archive contains:

- both raw evaluation JSON files;
- a machine-readable audit summary;
- the deterministic audit script;
- the full bilingual technical report;
- SHA-256 checksums.

No reconstruction source code, executable, or submission archive is included in the public attachment. The reproducible candidate submission can be provided privately to maintainers if needed.

[programbench_csview_public_issue_evidence_v1.zip](https://github.com/user-attachments/files/30601982/programbench_csview_public_issue_evidence_v1.zip)

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.