allenai / allenai/discoverybench
Real dataset and Reflexion
- Lingua principale
- Python
- Stelle
- 161
- Fork
- 18
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hi,
Thanks for creating this benchmarking dataset. This will be very helpful on building autonomous scientific discovery using LLMs.
I noticed the paper mentioned there were 264 tasks in the 'real' set but this repository has 283 'queries' in the metadata json files and the answer_key_real.csv has 239 rows. I'm wondering what in this repo is defined as the task in the paper. Can you give a hint on how to find those 264 tasks and the answers?
Besides, is it possible that you can share the prompt to run the Reflexion (oracle) method? Thanks!
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.