a2ui-project / a2ui-project/a2ui
[FEATURE]: Make the express format perform well with Gemma 2B models
- Lenguaje dominante
- TypeScript
- Estrellas
- 16.4k
- Forks
- 1.3k
- Merge medio
- 2 d 13 h
- PR fusionados (30 d)
- 134
Descripción
https://github.com/a2ui-project/a2ui/pull/2551 proposes a new "Vertical" inference format which is faster and more accurate, especially on Gemma models.
Vertical performs way better than Express. But, potentially we could update the Express system prompt to be more descriptive, or make the compiler more permissive, or change the format in some way to avoid the types of errors that we see.
Steps
- Use the new eval cases added in https://github.com/a2ui-project/a2ui/pull/2551 and evaluate the same Gemma 2B model with Express format
- Brainstorm different changes we could make to the express prompt or parser, and see if they fix the issues. We can just do small runs, e.g. 5-10 data points to verify this at first. Let's prefer simple tweaks e.g. to the prompt first, before making more radical changes the compiler or format itself.
- Prepare a report on the different approaches that are possible, and how effective they are. They could be implemented as "options" on the express format perhaps. At this point, consider doing larger eval runs with 2-5 runs per data point to build confidence that the fixes work well.
Guía de contribución
Línea de trabajo
Review the new eval cases from PR #2551. Run the Gemma 2B model with the existing Express format to benchmark performance. Experiment with tweaks to the Express system prompt, then test on 5-10 data points. Document the effectiveness of each approach and consider making them configurable options.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- machine-learning, typescript
- Área
- ai, tooling
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Activo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 45/100