InternLM / InternLM/InternLM-Math

Model evaluation on minif2f fails?

Abierto
#22 18 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
550
Forks
39
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

When I evaluate InternLM2-Math-Plus-7b in minif2f through this code, it fails. The model only generates one line "Here is the predicted next tactic:" without any tactics. If I let the model continue generating until I get a tactic each time, I only get a pass rate 38.1% instead of 43.4%.

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Línea de trabajo

Start by reproducing the InternLM2-Math-Plus-7b evaluation through the repository's minif2f evaluation code and inspect why generation stops after "Here is the predicted next tactic:". Compare the behavior with the reported 43.4% and 38.1% pass rates; done means the failure is explained and the evaluation produces the intended tactic output or a documented limitation.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
python
Área
machine-learning, testing-qa
Tipo de issue
Error
Dificultad
4/5
Tiempo estimado
3-5 días
Estado de actividad
Estancado
Claridad
Necesita aclaración
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.