deepspeedai / deepspeedai/DeepSpeed

[BUG] ZenFlow: loss values becomes NaN after 'update_interval' number of steps.

Open
#7,759 3 comments 0 reactions 1 assignee View on GitHub

@Antlera is already working on this.

Since Jan 4, 2026.

bug training
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

I was training a Llama model using the script released in the Deepspeed-Examples repository using ZenFlow. Interestingly, I noticed the loss values becoming NaN after the preconfigured (in the deepspeed config) update_interval number of steps. Ex: the update interval is set to 4 in the deepspeed config, loss becomes nan from the 5th step. The code used is available here: https://github.com/deepspeedai/DeepSpeedExamples/tree/master/training/DeepSpeed-ZenFlow/finetuning

My software configurations and versions:

torch==2.5.0+cu118
Transformers==4.57.3
Cloned the latest version of deepspeed from github
Datasets==4.4.1

The jobs are run with DGX H200 GPUs and AMD EPYC 7742 processor.

I wish to understand why the loss turns NaN and more importantly why it happens only after the update_interval step. Can I please get some help with respect to this ?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.