ruvnet / ruvnet/RuVector

[PIR][WP25] AI4AI-Bench (arXiv:2608.20318) recursive-improvement benchmark adapter (ADR-328)

Open
#894 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

adr phase-w4-1 pir wave-4
Dominant language
Rust
Stars
4.5k
Forks
603
Avg merge
23h 32m
Merged PRs (30d)
59

Description

Part of PIR Wave 4 — epic #837.

Add an AI4AI-Bench adapter behind the harness's constructor-injected benchmark seam (RuvectorFlywheelOptions.benchmark / RuvectorGepaOptions.benchmark / runControlledBenchmark runner param), adapting the released Apache-2.0 suite (github.com/Einsia/AI4AI-Bench). Contained-subprocess execution only; suite pinning; the flywheel.ts research-artifact-emission restriction stays intact. Deliverable: adapter + contract tests + at least one smoke-scale task run (acceptance-test clause 1). Official-scale B300 runs are USER ACTION.

  • Evidence grade & artifact availability: docs/research/perpetual-intelligence-runtime/10-wave4-evidence-review.md (PR #892)
  • WP/ADR mapping, phases, and security gates: docs/research/perpetual-intelligence-runtime/11-wave4-program-plan.md (PR #892)
  • ADR: docs/adr/ADR-328-ai4ai-bench-recursive-improvement-adapter.md

Depends on: WP2 (promotion pipeline), WP21 (external-grounding veto registration).

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read docs/adr/ADR-328-ai4ai-bench-recursive-improvement-adapter.md and the Wave 4 plan first, then trace the benchmark seams in RuvectorFlywheelOptions, RuvectorGepaOptions, runControlledBenchmark, and flywheel.ts. Done means a pinned AI4AI-Bench adapter with contained-subprocess execution, contract tests, and at least one smoke-scale task run; official B300 runs remain user action.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, security, testing
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.