facebookresearch / facebookresearch/ProgramBench

Clarify and harden cleanroom rules around reference executable instrumentation

Offen
#44 1 Kommentar 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Vorherrschende Sprache
Python
Sterne
924
Forks
67
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

Hi ProgramBench team, thanks for releasing the benchmark.

I wanted to ask for clarification about the intended cleanroom boundary for the provided reference executable during inference.

The current ProgramBench / mini-SWE-agent instructions are very clear that agents should infer behavior only by running the provided executable and reading bundled documentation. The default ProgramBench config also disallows internet access (`--network none`), runs as a non-root user, and drops `SYS_PTRACE`. The prompt explicitly forbids source lookup, wrapping/reusing the original binary, decompilation/disassembly, and `strace`/`ltrace` or similar instrumentation.

My question is whether the following should also be explicitly considered cleanroom violations, which I have observed during inference of my tested model:

- executing the reference with a polluted environment, e.g. `PATH=/tmp:$PATH ./executable ...` to make it call agent-written fake dependencies
- using loader instrumentation such as `LD_PRELOAD`, `LD_LIBRARY_PATH`, or dynamic linker tricks against `./executable`
- inspecting runtime process state via `/proc/$pid/{maps,fd,environ,cmdline}` or core/memory dumps
- changing cwd/tmp/config files specifically to observe implementation-level effects rather than normal public CLI behavior

These are different from normal allowed black-box probing, such as running `./executable --help`, passing inputs, and observing stdout/stderr/exit codes/filesystem side effects.

The current recommended Docker setup appears to run the agent and the reference executable in the same container. This blocks important classes of abuse (`--network none`, non-root user, `SYS_PTRACE` dropped), but it does not fully isolate the true reference executable from agent-controlled environment variables, writable `/tmp`/workspace state, PATH/cwd pollution, or same-container process observation.

Would you consider either:

1. documenting the above behaviors explicitly as disallowed instrumentation/wrapping of the oracle, and/or
2. adding an optional hardened inference harness where the true reference executable runs behind a separate sanitized oracle process/container, with the agent only able to issue structured black-box execution requests?

This is not intended as a security vulnerability report. It is a benchmark semantics / reproducibility question: the goal is to ensure scores measure behavioral inference from the public interface, not implementation-level oracle instrumentation.

Beitragsleitfaden

Beitragsleitfaden öffnen

Rechercherichtung

Beginne mit den aktuellen Anweisungen für ProgramBench und mini-SWE-agent sowie dem im Issue beschriebenen empfohlenen Docker-Setup. Vergleiche das ausdrücklich erlaubte Black-Box-Probing mit der Instrumentierung von Umgebung, Loader, Prozesszustand und Dateisystem und bestimme anschließend, ob die dokumentierte Cleanroom-Grenze ausreicht oder ein isoliertes strukturiertes Oracle-Harness erforderlich ist. Als erledigt gilt die Aufgabe, wenn die Regeln und, falls verfolgt, das Verhalten des Harness explizit und reproduzierbar sind.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
docker
Bereich
documentation, infrastructure, security
Issue-Typ
Feature
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Ruhig
Klarheit
Größtenteils klar
Anfängerfreundlichkeit
45/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.