InternLM / InternLM/InternLM-Math

Model evaluation on minif2f fails?

Open
#22 18 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
550
Forks
39
PR merge metrics
No merged PRs in 30d

Description

When I evaluate InternLM2-Math-Plus-7b in minif2f through this code, it fails. The model only generates one line "Here is the predicted next tactic:" without any tactics. If I let the model continue generating until I get a tactic each time, I only get a pass rate 38.1% instead of 43.4%.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reproducing the InternLM2-Math-Plus-7b evaluation through the repository's minif2f evaluation code and inspect why generation stops after "Here is the predicted next tactic:". Compare the behavior with the reported 43.4% and 38.1% pass rates; done means the failure is explained and the evaluation produces the intended tactic output or a documented limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.