LLM Eval UI work
- Dominant language
- Rust
- Stars
- 40
- Forks
- 3
- PR merge metrics
- No merged PRs in 30d
Description
### Description
Things that need to get done
- [ ] GenAIEvalDataSet registration
- [ ] Eval experimentation UI work
- [ ] Automated regression testing
Allow user to define (1) input and output schema of service, (2) build tasks to associate with eval record or use those defined in profile, (3) execute service feeding defined inputs and update eval context with output fields. Need to tag records so that we can pull observability metrics
Contributor guide
Research direction
No files, tests, or entry points are named. Start by locating GenAIEvalDataSet registration and the existing evaluation UI, then determine how schemas, profile tasks, service execution, eval context outputs, and observability tags are represented. Done means the UI supports the listed workflow and automated regression coverage exists.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- rust
- Domain
- ai, frontend, machine-learning, observability, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100