kohya-ss / kohya-ss/sd-scripts

Kaggle FP16 Stable Diffusion XL (SDXL) DreamBooth is being extremely inferior to BF16 supporting GPU training any ideas why could be?

Open
#1,031 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
7.2k
Forks
1.2k
Avg merge
11m
Merged PRs (30d)
2

Description

On Kaggle it does dual T4 GPU training. I used the very latest version with --ddp_gradient_as_bucket_view

Any ideas what could be causing so massive difference?

This was always the case since last few months whenever I did training

Everything is same except on Kaggle FP16 and xFormers enabled

SDXL 1.0 base DreamBooth training

![grid-0006](https://github.com/kohya-ss/sd-scripts/assets/19240467/75ecfdfe-2b63-4d11-bf38-e7d9612b9fc1)

![grid-0007](https://github.com/kohya-ss/sd-scripts/assets/19240467/081900d4-fc92-48b6-ae71-b6761d4f74e2)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by comparing the dual-T4 Kaggle run with the BF16-supporting GPU run, focusing on the stated FP16, xFormers, --ddp_gradient_as_bucket_view, and SDXL 1.0 DreamBooth settings. Reproduce the difference while changing one setting at a time; done means the cause of the inferior training result is identified and documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.