kohya-ss / kohya-ss/sd-scripts

Text encoder of SD 1.5 model is not trained which is not supposed to happen

Open
#855 52 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
7.2k
Forks
1.2k
Avg merge
11m
Merged PRs (30d)
2

Description

Here the executed command

accelerate launch --num_cpu_threads_per_process=2 "./train_db.py" --pretrained_model_name_or_path="/workspace/stable-diffusion-webui/models/Stable-diffusion/Realistic_Vision_V5.1.safetensors" --train_data_dir="/workspace/stable-diffusion-webui/models/Stable-diffusion/img" --reg_data_dir="/workspace/stable-diffusion-webui/models/Stable-diffusion/reg" --resolution="768,768" --output_dir="/workspace/stable-diffusion-webui/models/Stable-diffusion/model" --logging_dir="/workspace/stable-diffusion-webui/models/Stable-diffusion/log" --save_model_as=safetensors --full_bf16 --output_name="me_1e7" --lr_scheduler_num_cycles="4" --max_data_loader_n_workers="0" --learning_rate="1e-07" --lr_scheduler="constant" --train_batch_size="1" --max_train_steps="4160" --save_every_n_epochs="1" --mixed_precision="bf16" --save_precision="bf16" --cache_latents --cache_latents_to_disk --optimizer_type="Adafactor" --optimizer_args scale_parameter=False relative_step=False warmup_init=False weight_decay=0.01 --max_data_loader_n_workers="0" --bucket_reso_steps=64 --xformers --bucket_no_upscale --noise_offset=0.0

When text encoder is not trained it is supposed to print `Text Encoder is not trained.`

This message is not printed either

```
train_text_encoder = args.stop_text_encoder_training is None or args.stop_text_encoder_training >= 0
unet.requires_grad_(True) # 念のため追加
text_encoder.requires_grad_(train_text_encoder)
if not train_text_encoder:
accelerator.print("Text Encoder is not trained.")
```

So how do I know text encoder were not trained? Because I extracted LoRA and it says text encoder is same

I did 30 trainings and so many trainings are wasted because of this bug :/

![image](https://github.com/kohya-ss/sd-scripts/assets/19240467/ca373092-cd25-4d16-b256-4bb24eec9e78)

@kohya-ss

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the executed command and inspect train_db.py around the shown train_text_encoder and accelerator.print logic. Reproduce the reported SD 1.5 training case, then verify whether the text encoder training state and status message agree with the extracted LoRA result; done means the behavior is accurately reported.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.