awslabs / awslabs/agent-evaluation

LLM hallucinates when converting the 'steps' into actual prompts for the Evaluator

Open
#107 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
372
Forks
51
PR merge metrics
No merged PRs in 30d

Description

**My test file says this:**
"Give me a competitive OP01-060 leader decklist from a recent tournament."

**The LLM gets turned into this:**
"I'm looking for a competitive OP01-060 (Operational Program Duelist Alliance) leader decklist from any recent major Yu-Gi-Oh! tournament, regardless of region."

It is a totally different prompt from the specified step. Is there anyway to tune the LLM to not be so creative?

Contributor guide

Open the contributing guide

Research direction

No source file, test, or entry point is named. Start by locating the evaluator path that converts written steps into prompts and reproduce the OP01-060 example; done means the generated prompt preserves the specified step without adding or changing its meaning.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.