aws-samples / aws-samples/remote-swe-agents

Auto-evaluation of agent behaviors

Open
#44 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
TypeScript
Stars
243
Forks
50
Avg merge
1h 19m
Merged PRs (30d)
24

Description

We want to introduce some evaluation framework like promptfoo to automatically evaluate the agents performance and to be more confident on each change to prompts or code.

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named. Start by surveying the TypeScript agent and prompt execution paths, then review how changes are currently checked. Done means the repository has an evaluation framework that automatically assesses agent behavior after prompt or code changes.

Written by the indexing model from the issue text.

Assessment

Tech stack
typescript
Domain
ai, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.