a2ui-project / a2ui-project/a2ui
Central A2UI evaluation harness
- Dominant language
- TypeScript
- Stars
- 16.4k
- Forks
- 1.3k
- Avg merge
- 3d 15h
- Merged PRs (30d)
- 134
Description
Central A2UI evaluation harness and data points covering top use cases and models, to allow us to make further data-driven optimizations. See [A2UI Inference Performance Project Scope](https://docs.google.com/document/d/1BDILK5Wktik1Pce_pjPBdRoVTWWJQSx9y7gJvObwy4c/edit?tab=t.0) (internal Google document)
Contributor guide
Research direction
The issue references an internal Google document for scope, which is not accessible. Without the document, the specific requirements, entry points, or existing evaluation code are unknown. A newcomer would need to understand the A2UI project's architecture, existing inference pipelines, and define what metrics and models to evaluate. Start by looking for existing evaluation scripts or benchmarks in the repository.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai-infra-agents, tooling
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100