modelscope / modelscope/DiffSynth-Studio
Expected to have finished reduction in the prior iteration before starting a new one.
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 13.1k
- Forks
- 1.3k
- Avg merge
- 13h 12m
- Merged PRs (30d)
- 45
Description
I was finetuning the Flux.1-dev: Upscaler ControlNet model with my custom data, but met this problem during training:
[rank1]: RuntimeError: Expected to have finished reduction in the prior iteration before starting a new one. This error indicates that your module has parameters that were not used in producing loss. You can enable unused parameter detection by passing the keyword argument find_unused_parameters=Truetotorch.nn.parallel.DistributedDataParallel``
Here are my training script:
#! /bin/bash
export NUM_NODES=1
export NUM_GPUS=2
export NCCL_DEBUG=IVFO
export CUDA_LAUNCH_BLOCKING=1
DATASET_BASE_PATH="/workspace/codes/DiffSynth-Studio/data/neemo_mini_1440p_120f/controlnet_data"
DATASET_METADATA_PATH="${DATASET_BASE_PATH}/metadata.csv"
OUTPUT_PATH="/workspace/codes/DiffSynth-Studio/outputs/neemo_mini_1440p_120f/train_finetune_controlnet/models/FLUX.1-dev-Controlnet-Upscaler_lora"
IMG_HEIGHT=1440
IMG_WIDTH=2560
DATASET_REPEAT=10
NUM_EPOCHS=5
accelerate launch --mixed_precision=bf16 --multi_gpu --main_process_port 29501 --num_machines $NUM_NODES --num_processes $NUM_GPUS --config_file examples/flux/model_training/full/accelerate_config.yaml examples/flux/model_training/train.py \
--dataset_base_path $DATASET_BASE_PATH \
--dataset_metadata_path $DATASET_METADATA_PATH \
--data_file_keys "image,controlnet_image" \
--dataset_repeat $DATASET_REPEAT \
--height $IMG_HEIGHT \
--width $IMG_WIDTH \
--model_id_with_origin_paths "black-forest-labs/FLUX.1-dev:flux1-dev.safetensors,black-forest-labs/FLUX.1-dev:text_encoder/model.safetensors,black-forest-labs/FLUX.1-dev:text_encoder_2/,black-forest-labs/FLUX.1-dev:ae.safetensors,jasperai/Flux.1-dev-Controlnet-Upscaler:diffusion_pytorch_model.safetensors" \
--learning_rate 1e-5 \
--num_epochs $NUM_EPOCHS \
--remove_prefix_in_ckpt "pipe.controlnet.models.0." \
--output_path $OUTPUT_PATH \
--trainable_models "controlnet" \
--extra_inputs "controlnet_image" \
--use_gradient_checkpointing
Below is the complete logs:
job-f2c0d17be585-20251017213424-worker-0.log
I am appreciate to any solution or suggestions. 🙏
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with examples/flux/model_training/train.py and examples/flux/model_training/full/accelerate_config.yaml, then reproduce the provided two-GPU command using the linked worker log. Trace the DistributedDataParallel setup and the trainable controlnet path; done means identifying a reproducible cause and documenting or validating a project-level fix.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch, yaml
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100