aws / aws/fmeval

[Feature] LLM-based (QA Accuracy) eval algorithm

Open
#163 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
291
Forks
60
PR merge metrics
No merged PRs in 30d

Description

The metrics-based approaches in the `QAAccuracy` eval algorithm seem to harshly penalize verbose models (like Claude) on datasets with concise reference answers (like SQuAD).

It'd be useful if this library could provide support for LLM-based evaluation of LLM results: For example asking a model whether the reference answer and the generated answer agree or disagree. I'd imagine it working something along the lines of [LlamaIndex's `CorrectnessEvaluator`](https://github.com/run-llama/llama_index/blob/06b0e390fe1d2fab7c41ff63f6505810c15f8684/llama_index/evaluation/correctness.py)?

As I understand it should be possible in theory to implement something like this by building a custom `EvalAlgorithmInterface`-based class, but there are a lot of design questions to consider like:

- Is it possible to control whether the same LLM, a different LLM, or a panel of multiple LLMs gets used for the evaluation step, versus the original answer generation?
- Since there are lots of different ways to use LLMs for self-critique, maybe e.g. `QAAccuracyByLLMCritic` should be a subtype of some broader class? Certainly it'd be interesting to use LLMs to judge other aspects like relevancy, and specific aspects of tone e.g. ("did it discuss my competitor companies XYZ")

Contributor guide

Open the contributing guide

Research direction

Start by reading the QAAccuracy eval algorithm and the EvalAlgorithmInterface, then compare the linked LlamaIndex CorrectnessEvaluator. The issue names evaluator-model selection, multi-model judging, and possible broader abstractions as unresolved design questions. Done would require an agreed API and implementation for LLM-based QA evaluation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.