AlibabaResearch / AlibabaResearch/DAMO-ConvAI

Evaluation Method for BIRD Dataset [Enhancement]

Aperta
#159 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
1.6k
Fork
250
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

Hello,

I encountered an improvement opportunity during the evaluation process for the BIRD dataset. The prediction below is marked as incorrect by the evaluation method, but the only difference is the order of the elements.

![image](https://github.com/user-attachments/assets/f706bc94-fd96-483b-8e5a-9b9739e0c5ae)

The evaluation method uses strict equality. This occurs in the file [bird/llm/src/evaluation.py](https://github.com/AlibabaResearch/DAMO-ConvAI/blob/main/bird/llm/src/evaluation.py) on line 26.

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.