AI-Hypercomputer / AI-Hypercomputer/maxtext

Issue with gradient accumulation and expansion_factor_real_data>1

Abierto
#3,381 0 comentarios 0 reacciones 1 asignado Reclamado por @khatwanimohit Ver en GitHub
bug
Lenguaje dominante
Python
Estrellas
2.4k
Forks
607
Merge medio
2 d 19 h
PR fusionados (30 d)
158

Descripción

### Bug report

`expansion_factor_real_data` > 1 generates placeholder data and then truncates the inputs in `loss_fn` to the correct size, which is supposed to eliminate the placeholder data. In gradient accumulation, from `train_step` we reshape the inputs to introduce a `gradient_accumulation_steps` dimension.

If we first truncated then reshaped, this would work correctly. However, we reshape then truncate, which means later gradient accumulation steps use the placeholder data.

I believe `max_checkify` does not catch this issue because it happens too early in the process.

(Internally we're on an older fork of this codebase, so I apologize if this has been fixed already. I looked through the relevant code and it looked like it would have the same issue)

### Logs/Output

_No response_

### Environment Information

_No response_

### Additional Context

_No response_

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.