facebookresearch / facebookresearch/ProgramBench
`eval_clean_hashes` is missing or stale for some `task_cleanroom_v6` images
- Dominant language
- Python
- Stars
- 924
- Forks
- 67
- PR merge metrics
- No merged PRs in 30d
Description
## Summary
`eval_clean_hashes` is intended to remove byte-for-byte copies of the gold executable from a submission before `compile.sh` runs. However, some task metadata does not define the field, and some existing values do not match the reference executable currently shipped in the corresponding `task_cleanroom_v6` image.
This makes the metadata incomplete for downstream evaluators that rely on `eval_clean_hashes` to scrub renamed copies of the reference executable.
## Current state
At repository commit `963063c`:
- 201 task directories total (including `testorg__calculator.abc1234`)
- 40 task YAML files do not define `eval_clean_hashes`
- 161 define at least one hash
- none define an explicitly empty list
Excluding the test fixture, 39 of the 200 real tasks are missing the field.
In addition, a non-empty list does not necessarily contain the SHA-256 of the executable in the current `task_cleanroom_v6` image.
Examples observed:
| Task | Hash in `task.yaml` | Current `task_cleanroom_v6` reference SHA-256 |
|---|---|---|
| `facebookresearch__fasttext.1142dc4` | missing | `80310c0ae4d92165b7a94b17439a1931d99567ae4c1d54706bb557e87bdcd51c` |
| `halitechallenge__halite.822cfb6` | missing | `9d5e3ec763bbfa7c1091054c3a746ee25af7d4368edc4395eb3e3d8f7f9feac3` |
| `tomnomnom__gron.88a6234` | missing | `0270943958c042821654a1e19608848a8cf14cd1136fde4726e8a1821cdf4cc7` |
| `rs__curlie.5dfcbb1` | `e7e4f84ba194306ef782853d8b77788dd4674a6baad19c878885bb08c2db0d35` | `7ebf850e7d8d0def143b72a1d6d0ebf4e54d8212c515c2317f06e4e8f3f642f0` |
| `wfxr__csview.8ac4de0` | `e07420f640d1e6a0c047cab49431d1b561f661241efe111a053ff5dde1119d09` | `758ef03ba091d6a7c77d32e60aaf484aca23532aba3eea6a51bdf25232c5f5ce` |
As a control, `xorg62__tty-clock.f2f847c` does contain the current reference hash (`cd400708...`).
## Reproduction
For a task image:
```bash
docker run --rm \
--network none \
--user root \
--entrypoint sha256sum \
programbench/rs_1776_curlie.5dfcbb1:task_cleanroom_v6 \
/workspace/executable
```
Compare the resulting digest with:
```text
src/programbench/data/tasks/rs__curlie.5dfcbb1/task.yaml
```
The evaluator consumes this metadata in `eval_batch.py` and recursively removes matching files in `Evaluator._remove_hashed_files()` before compilation. The evaluator comment describes these hashes as protection against byte-for-byte copies of the gold binary, so including the current cleanroom reference hash appears consistent with the existing field semantics.
## Suggested fix
1. Compute the SHA-256 of `/workspace/executable` for every current `task_cleanroom_v6` image.
2. Append it to that task's `eval_clean_hashes`, preserving historical hashes and deduplicating the list.
3. Add a maintenance command or release script that updates these values whenever cleanroom images are rebuilt.
4. Validate that every real task has at least one well-formed 64-character SHA-256 value.
5. Ideally pin cleanroom images by digest, or verify/update the hashes as part of the image publication workflow, so mutable tags cannot silently drift from task metadata.
I can prepare a follow-up PR to populate the current hashes if this approach matches the intended use of `eval_clean_hashes`.
Contributor guide
Assessment
This issue has not been assessed yet.