Reproducing existing results on NarrativeQA
未關閉
- 主要語言
- Python
- 星號
- 2.4k
- 分支
- 201
- PR 合併指標
- 30 天內沒有已合併 PR
描述
I'm trying to reproduce the results for NarrativeQA by directly running the command with the .yml configuration files. Below are the performances measured with ROUGE-L-Max.
For PPO with supervision, I got 0.581 and 0.588 for epochs 0 and 99, respectively.
For NLPO with supervision, I got **0.217** and **0.213** for epochs 0 and 99, respectively.
I'm wondering why the result for NLPO doesn't match the reported result in the paper.
I also tried to use the config for PPO, and just modify the RL algorithm to NLPO, I got the same result as above.
Please let me know if I'm missing something or if it's some other issue. Thanks!
貢獻指南
這個儲存庫沒有索引到貢獻指南
評估
這個 Issue 還沒有評估資料。