huggingface / huggingface/evaluate

Feature Request: Add support for retrieving mispredicted examples with predictions and ground truth

Open
#672 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
2.5k
Forks
341
PR merge metrics
No merged PRs in 30d

Description

First of all, thank you for the great work on the evaluate library — it's been incredibly useful for benchmarking and analyzing model performance!

I’d like to suggest a feature idea that could enhance the interpretability and debugging process when evaluating models. Specifically, it would be great if evaluate could provide a way to retrieve the list of mispredicted examples, including both the model's prediction and the corresponding ground truth.

This could be especially useful for Creating visualizations or reports of incorrect predictions

Thanks again 🙌

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.