InternLM / InternLM/InternLM-Math
Model evaluation on minif2f fails?
- 主要言語
- Python
- スター
- 550
- フォーク
- 39
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
When I evaluate InternLM2-Math-Plus-7b in minif2f through this code, it fails. The model only generates one line "Here is the predicted next tactic:" without any tactics. If I let the model continue generating until I get a tactic each time, I only get a pass rate 38.1% instead of 43.4%.
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
Start by reproducing the InternLM2-Math-Plus-7b evaluation through the repository's minif2f evaluation code and inspect why generation stops after "Here is the predicted next tactic:". Compare the behavior with the reported 43.4% and 38.1% pass rates; done means the failure is explained and the evaluation produces the intended tactic output or a documented limitation.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- machine-learning, testing-qa
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 25/100