allenai / allenai/discoverybench

[EVALUATION]: How to evaluate on the test dataset absent gold hypothesis?

未關閉
#20 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
evaluation
主要語言
Python
星號
161
分支
18
PR 合併指標
30 天內沒有已合併 PR

描述

The `meatadata_*.json` files under `discoverybench/real/test` do not seem to contain labeled hypothesis. How should we evaluate this portion of the dataset to get HMS scores?

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。