allenai / allenai/fluid-benchmarking

Issue with reproducing the paper results

オープン
#2 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
29
フォーク
4
PR マージ指標
30日以内にマージされた PR はありません

説明

Hello,

Could you please provide the evaluation metrics codes to produce Tables 1 and 2 ? After all my attempts I could not reproduce those numbers (I did get the main trend but the numbers are far off good as reported in the paper). Even when using the experiments.jsonl provided in the repository (and the one obtained from running scripts/run_experiments.py) I still can not get the numbers stated in the paper. Also, the both files contain 2,712 ckpts, not 2802 as mentioned on the paper.

Any help would be greatly appreciated. Thanks in advance

CC @valentinhofmann @davidheineman

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。