AlibabaResearch / AlibabaResearch/DAMO-ConvAI
Evaluation Method for BIRD Dataset [Enhancement]
Aperta
- Lingua principale
- Python
- Stelle
- 1.6k
- Fork
- 250
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hello,
I encountered an improvement opportunity during the evaluation process for the BIRD dataset. The prediction below is marked as incorrect by the evaluation method, but the only difference is the order of the elements.

The evaluation method uses strict equality. This occurs in the file [bird/llm/src/evaluation.py](https://github.com/AlibabaResearch/DAMO-ConvAI/blob/main/bird/llm/src/evaluation.py) on line 26.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.