AI-Hypercomputer / AI-Hypercomputer/maxtext
Why logit checker has such a high tolerance?
- Lenguaje dominante
- Python
- Estrellas
- 2.4k
- Forks
- 607
- Merge medio
- 2 d 19 h
- PR fusionados (30 d)
- 158
Descripción
Hi, I would like to understand why the comparison with the golden logits uses high absolute tolerance values, such as:
- `--atol=3` for [Mistral-7b](https://github.com/AI-Hypercomputer/maxtext/blob/1ff8505cb2c9fb5abd65de73cf3d9b70cec20890/end_to_end/tpu/mistral/7b/test_mistral-7b.sh#L42)
- `--atol=1.0` for [Gemma2-9b](https://github.com/AI-Hypercomputer/maxtext/blob/1ff8505cb2c9fb5abd65de73cf3d9b70cec20890/end_to_end/tpu/gemma2/9b/2_test_gemma.sh#L47)
while some models achieve more similar results with smaller tolerances:
- `--atol=0.2` for [Llama-2-70b](https://github.com/AI-Hypercomputer/maxtext/blob/1ff8505cb2c9fb5abd65de73cf3d9b70cec20890/end_to_end/tpu/llama2/70b/2_test_llama2_70b.sh#L63C53-L63C63)
Why is this high tolerance needed? I understand we might expect numerical differences, but I was hoping for more similar results when implementations are equivalent.
Guía de contribución
Evaluación
Este issue todavía no se ha evaluado.