AI-Hypercomputer / AI-Hypercomputer/maxtext

Apparent bug in the megablocks implementation

未关闭
#1,183 14 条评论 2 个 reaction 已指派 1 人 已被 @RissyRan 认领 在 GitHub 查看
主要语言
Python
星标
2.4k
派生
607
平均合并
2 天 19 小时
30 天内合并 PR
158

描述

Hello,

Firstly, thank you very much for providing us with a great industry-grade LLM training library.

I've noticed that when `megablox=True`, the logits do not match those of the Huggingface implementation: [link to the specific code](https://github.com/AI-Hypercomputer/maxtext/blob/main/end_to_end/tpu/mixtral/8x7b/2_test_mixtral.sh#L46).

Additionally, when fine-tuning from the mixtral checkpoint, the loss begins higher than expected but rapidly decreases. However, the resulting model weights, when converted back to the Huggingface format, perform poorly on MMLU.

Conversely, when `sparse_matmul=True` and `megablox=False`, the loss starts at a lower level and the resulting Huggingface-converted model performs well on MMLU. Nevertheless, the MFU is approximately 3 times lower with `ragged_dot` than with `megablox`, making training impractical at larger scales.

Are there any plans to address these discrepancies in the implementation?

Best regards.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。