allenai / allenai/longformer

fail to reproduce the base model result of the TriviaQA Dataset with scripts/trivia.py.

オープン
#138 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Python
スター
2.2k
フォーク
286
PR マージ指標
30日以内にマージされた PR はありません

説明

hi,
The f1score in the paper is 75.2 but I only got 74.7 by using trivia.py.
And I have a few questions.

1. The triviaQA attention window hyperparameter is not given in the paper, and in the longformer-base-406/config.json, the attention window is all set to 256, Is this correct?

2. The epoch in the paper is set to 5, but the cheatsheet.txt is set to 4.

3. find some bug in trivia.py.
https://github.com/allenai/longformer/blob/0674e0e7bf10007e0dafd7fb65befe96c11bfcb6/scripts/triviaqa.py#L678
num_devices = 1 or len(args.gpus).
In python, 1 or 8 will be equal to 1.

Thanks.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。