kohya-ss / kohya-ss/sd-scripts
Kaggle FP16 Stable Diffusion XL (SDXL) DreamBooth is being extremely inferior to BF16 supporting GPU training any ideas why could be?
- Dominant language
- Python
- Stars
- 7.2k
- Forks
- 1.2k
- Avg merge
- 11m
- Merged PRs (30d)
- 2
Description
On Kaggle it does dual T4 GPU training. I used the very latest version with --ddp_gradient_as_bucket_view
Any ideas what could be causing so massive difference?
This was always the case since last few months whenever I did training
Everything is same except on Kaggle FP16 and xFormers enabled
SDXL 1.0 base DreamBooth training


Contributor guide
No contributing guide indexed for this repository
Research direction
Start by comparing the dual-T4 Kaggle run with the BF16-supporting GPU run, focusing on the stated FP16, xFormers, --ddp_gradient_as_bucket_view, and SDXL 1.0 DreamBooth settings. Reproduce the difference while changing one setting at a time; done means the cause of the inferior training result is identified and documented.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100