aws-samples / aws-samples/remote-swe-agents
Auto-evaluation of agent behaviors
Open
- Dominant language
- TypeScript
- Stars
- 243
- Forks
- 50
- Avg merge
- 1h 19m
- Merged PRs (30d)
- 24
Description
We want to introduce some evaluation framework like promptfoo to automatically evaluate the agents performance and to be more confident on each change to prompts or code.
Contributor guide
Research direction
No files, tests, or entry points are named. Start by surveying the TypeScript agent and prompt execution paths, then review how changes are currently checked. Done means the repository has an evaluation framework that automatically assesses agent behavior after prompt or code changes.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- typescript
- Domain
- ai, testing-qa
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100