allenai / allenai/discoverybench
[EVALUATION]: How to evaluate on the test dataset absent gold hypothesis?
未關閉
evaluation
- 主要語言
- Python
- 星號
- 161
- 分支
- 18
- PR 合併指標
- 30 天內沒有已合併 PR
描述
The `meatadata_*.json` files under `discoverybench/real/test` do not seem to contain labeled hypothesis. How should we evaluate this portion of the dataset to get HMS scores?
貢獻指南
這個儲存庫沒有索引到貢獻指南
評估
這個 Issue 還沒有評估資料。