InternLM / InternLM/InternLM-Math
Model evaluation on minif2f fails?
- Vorherrschende Sprache
- Python
- Sterne
- 550
- Forks
- 39
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
When I evaluate InternLM2-Math-Plus-7b in minif2f through this code, it fails. The model only generates one line "Here is the predicted next tactic:" without any tactics. If I let the model continue generating until I get a tactic each time, I only get a pass rate 38.1% instead of 43.4%.
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Rechercherichtung
Start by reproducing the InternLM2-Math-Plus-7b evaluation through the repository's minif2f evaluation code and inspect why generation stops after "Here is the predicted next tactic:". Compare the behavior with the reported 43.4% and 38.1% pass rates; done means the failure is explained and the evaluation produces the intended tactic output or a documented limitation.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- machine-learning, testing-qa
- Issue-Typ
- Bug
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100