AI-Hypercomputer / AI-Hypercomputer/maxtext

Apparent bug in the megablocks implementation

Abierto
#1,183 14 comentarios 2 reacciones 1 asignado Reclamado por @RissyRan Ver en GitHub
Lenguaje dominante
Python
Estrellas
2.4k
Forks
607
Merge medio
2 d 19 h
PR fusionados (30 d)
158

Descripción

Hello,

Firstly, thank you very much for providing us with a great industry-grade LLM training library.

I've noticed that when `megablox=True`, the logits do not match those of the Huggingface implementation: [link to the specific code](https://github.com/AI-Hypercomputer/maxtext/blob/main/end_to_end/tpu/mixtral/8x7b/2_test_mixtral.sh#L46).

Additionally, when fine-tuning from the mixtral checkpoint, the loss begins higher than expected but rapidly decreases. However, the resulting model weights, when converted back to the Huggingface format, perform poorly on MMLU.

Conversely, when `sparse_matmul=True` and `megablox=False`, the loss starts at a lower level and the resulting Huggingface-converted model performs well on MMLU. Nevertheless, the MFU is approximately 3 times lower with `ragged_dot` than with `megablox`, making training impractical at larger scales.

Are there any plans to address these discrepancies in the implementation?

Best regards.

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.