AI-Hypercomputer / AI-Hypercomputer/maxtext

DEFAULT_MASK_VALUE causes gradient explosion and nan loss on deep models

未关闭
#614 2 条评论 2 个 reaction 已指派 1 人 已被 @anfals 认领 在 GitHub 查看
bug
主要语言
Python
星标
2.4k
派生
607
平均合并
2 天 19 小时
30 天内合并 PR
158

描述

I was training a llama model on GPU, with a custom embedding. It worked fine with 12 layers, dim 1024, seq length 256, but loss would become nan after the first step if setting num_layers to more than 17. I debugged the gradients, and found after each layer their magnitude would increase by around 100x, until they hit float32_max at around the 18th layer and became inf, leading to nan loss.

The gradient explosion seemed to be coming from
`local_exps = jnp.exp(attn_weights - local_max)`
in attentions.py.

Changing

`DEFAULT_MASK_VALUE = -0.7 * float(jnp.finfo(jnp.dtype("float32")).max)`
to
`DEFAULT_MASK_VALUE = -jnp.inf`
fixed the issue, and the gradients' magnitude stopped increasing after each level.

Presumably the issue wasn't noticed during TPU training as that uses a separate codepath.

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。