[PIR][WP25] AI4AI-Bench (arXiv:2608.20318) recursive-improvement benchmark adapter (ADR-328)
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 4.5k
- Forks
- 603
- Avg merge
- 23h 32m
- Merged PRs (30d)
- 59
Description
Part of PIR Wave 4 — epic #837.
Add an AI4AI-Bench adapter behind the harness's constructor-injected benchmark seam (RuvectorFlywheelOptions.benchmark / RuvectorGepaOptions.benchmark / runControlledBenchmark runner param), adapting the released Apache-2.0 suite (github.com/Einsia/AI4AI-Bench). Contained-subprocess execution only; suite pinning; the flywheel.ts research-artifact-emission restriction stays intact. Deliverable: adapter + contract tests + at least one smoke-scale task run (acceptance-test clause 1). Official-scale B300 runs are USER ACTION.
- Evidence grade & artifact availability:
docs/research/perpetual-intelligence-runtime/10-wave4-evidence-review.md(PR #892) - WP/ADR mapping, phases, and security gates:
docs/research/perpetual-intelligence-runtime/11-wave4-program-plan.md(PR #892) - ADR:
docs/adr/ADR-328-ai4ai-bench-recursive-improvement-adapter.md
Depends on: WP2 (promotion pipeline), WP21 (external-grounding veto registration).
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Read docs/adr/ADR-328-ai4ai-bench-recursive-improvement-adapter.md and the Wave 4 plan first, then trace the benchmark seams in RuvectorFlywheelOptions, RuvectorGepaOptions, runControlledBenchmark, and flywheel.ts. Done means a pinned AI4AI-Bench adapter with contained-subprocess execution, contract tests, and at least one smoke-scale task run; official B300 runs remain user action.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai, security, testing
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100