huggingface / huggingface/diffusers

Add aspect ratio bucketing to training scripts

Open
#7,908 11 comments 3 reactions 0 assignees View on GitHub
stale
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

**Is your feature request related to a problem? Please describe.**
When fine tuning SDXL, images are required to be a fixed size (1024x1024) which involves a lot of cropping that both takes time/resources, and often causes important parts of the image to get cropped out, which lowers model quality.

**Describe the solution you'd like.**
The ideal solution would be a simple option for user to enable aspect ratio bucketing (e.g. a command argument `--enable-bucketing`) that will let them train with multiple image sizes

Contributor guide

Open the contributing guide

Research direction

Start by locating the SDXL training scripts and the code responsible for enforcing fixed image sizes and cropping. Define how the optional --enable-bucketing argument should select multiple image sizes, then verify that training can preserve varied aspect ratios without requiring every image to be cropped to 1024x1024.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.