allenai / allenai/longformer

fail to reproduce the base model result of the TriviaQA Dataset with scripts/trivia.py.

Ouverte
#138 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Python
Étoiles
2.2k
Forks
286
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

hi,
The f1score in the paper is 75.2 but I only got 74.7 by using trivia.py.
And I have a few questions.

1. The triviaQA attention window hyperparameter is not given in the paper, and in the longformer-base-406/config.json, the attention window is all set to 256, Is this correct?

2. The epoch in the paper is set to 5, but the cheatsheet.txt is set to 4.

3. find some bug in trivia.py.
https://github.com/allenai/longformer/blob/0674e0e7bf10007e0dafd7fb65befe96c11bfcb6/scripts/triviaqa.py#L678
num_devices = 1 or len(args.gpus).
In python, 1 or 8 will be equal to 1.

Thanks.

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.