huggingface / huggingface/diffusers

Support for Multi-Image Input LoRA Training Pipeline for FLUX.2-Klein (e.g., Style Transfer and Multi-Subject Composition)

Open
#13,008 2 comments 2 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
34.5k
Forks
7.3k
Avg merge
3d 3h
Merged PRs (30d)
91

Description

### Model/Pipeline/Scheduler description

I would like to request official support in Diffusers for a multi-image input LoRA training pipeline targeting the FLUX.2 Klein model. It seems that existing LoRA training pipelines are designed around single-image conditioning, which limits their applicability for tasks that naturally require multiple reference images. Clear guidance, reference implementations, or examples demonstrating how multi-image conditioning could be handled during training would be highly valuable. Thank you for your continued work on Diffusers, and I would greatly appreciate any insights, recommendations, or future plans related to supporting this capability.

### Open source status

- [ ] The model implementation is available.
- [ ] The model weights are available (Only relevant if addition is not a scheduler).

### Provide useful links for the implementation

_No response_

Contributor guide

Open the contributing guide

Research direction

No files or tests are named. Start by reviewing Diffusers' existing LoRA training pipelines and the FLUX.2 Klein model implementation and weights, then compare their single-image conditioning with the requested multi-image cases. Done means an agreed implementation scope plus official support or documented reference examples for style transfer and multi-subject composition.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
ai, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.