Azure / Azure/azureml-examples

rag evaluation: 'mean_gpt_retrieval_score': nan}

Open
#3,226 0 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Jupyter Notebook
Stars
2k
Forks
1.7k
Avg merge
18h 18m
Merged PRs (30d)
2

Description

### Operating System

MacOS

### Version Information

(modeleval) (base) MengMacBook-M3MaxPro:model_evaluation wanmeng$ pip install azureml-metrics[generative-ai]
Requirement already satisfied: azureml-metrics[generative-ai] in /Users/wanmeng/miniconda3/envs/modeleval/lib/python3.11/site-packages (0.0.57)

### Steps to reproduce

1. OPENAI_API_VERSION="2024-02-01"
OPENAI_API_BASE="https://gpt4o-eliz-westus3.openai.azure.com/"
OPENAI_API_TYPE="azure"
OPENAI_API_KEY="
deployment_id="gpt-4o"
2. %%time
# gpt4o model
from azureml.metrics import compute_metrics, constants
from pprint import pprint
import os

y_test = [["4", "2 + 2 = 4"], ["Agra", "Agra, India"]]

y_pred = [
[
{"role": "user", "content": "What is the value of 2 + 2?"},
{"role": "assistant", "content": "2 + 2 = 4",
"context": {
"citations": [{'id': 'math_document1.md',
'content': 'Information about additions: ' \
'1 + 2 = 3, 2 + 2 = 4'}]
}
}
],
[
{"role": "user", "content": "Where is Taj Mahal located?"},
{"role": "assistant", "content": "Taj Mahal is located in Agra, India",
"context": {
"citations": [{'id': 'taj_mahal_document1.md',
'content': 'Taj Mahal is located in Agra, India ' \
'and is one of the seven wonders of the world.'}]
}
}
]
]

openai_params = {
"api_version": OPENAI_API_VERSION,
"api_base": OPENAI_API_BASE,
"api_type": OPENAI_API_TYPE,
"api_key" : OPENAI_API_KEY,
"deployment_id": deployment_id
}

metrics_config = {
"openai_params": openai_params,
"score_version": "v1",
"use_chat_completion_api": True,
# To compute RAG based metrics
"metrics": ["gpt_relevance", "gpt_groundedness", "gpt_retrieval_score"]
}

# The above metrics can even be computed by setting the task_type to RAG_EVALUATION
result = compute_metrics(task_type=constants.Tasks.CHAT_COMPLETION,
y_test=y_test,
y_pred=y_pred,
**metrics_config)
pprint(result)

### Expected behavior

'metrics': {'mean_gpt_groundedness': 5.0,
'mean_gpt_relevance': 5.0,
'mean_gpt_retrieval_score': 5.0}}

### Actual behavior

'metrics': {'mean_gpt_groundedness': 5.0,
'mean_gpt_relevance': 5.0,
'mean_gpt_retrieval_score': nan}}

### Addition information

A customer is using this metrics next week, pls help fix the problem, thanks a lot!

Contributor guide

Open the contributing guide

Research direction

Reproduce the result with the provided Python chat-completion example and the azureml-metrics configuration. Trace compute_metrics and the gpt_retrieval_score metric to determine why the two supplied citations produce nan. Done means the same inputs return a numeric retrieval score matching the expected metrics output.

Written by the indexing model from the issue text.

Assessment

Tech stack
azure, jupyter-notebook, python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.