awslabs / awslabs/agent-evaluation
LLM hallucinates when converting the 'steps' into actual prompts for the Evaluator
- Dominant language
- Python
- Stars
- 372
- Forks
- 51
- PR merge metrics
- No merged PRs in 30d
Description
**My test file says this:**
"Give me a competitive OP01-060 leader decklist from a recent tournament."
**The LLM gets turned into this:**
"I'm looking for a competitive OP01-060 (Operational Program Duelist Alliance) leader decklist from any recent major Yu-Gi-Oh! tournament, regardless of region."
It is a totally different prompt from the specified step. Is there anyway to tune the LLM to not be so creative?
Contributor guide
Research direction
No source file, test, or entry point is named. Start by locating the evaluator path that converts written steps into prompts and reproduce the OP01-060 example; done means the generated prompt preserves the specified step without adding or changing its meaning.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100