Eval-ready run history and trajectory export
Open
enhancement
- Dominant language
- TypeScript
- Stars
- 354
- Forks
- 40
- PR merge metrics
- No merged PRs in 30d
Description
Persist complete, eval-ready run histories so past sessions can be graded, compared, and rerun against new models, prompts, skills, rubrics, and harness versions.
Contributor guide
Research direction
The issue names no files, tests, or entry points, so first map where runs and trajectories are currently created and stored. Clarify the export format and required metadata for models, prompts, skills, rubrics, and harness versions; done means a past session can be graded, compared, and rerun against those changing inputs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100