Very High VRAM usage when using lora with flux
Open
Potential Bug
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 7h
- Merged PRs (30d)
- 158
Description
### Expected Behavior
Not 10Gb Vram eaten using the lora.
### Actual Behavior
I have flux fp8 schnell on a 3090, I run two loras rank 64 onto the model, but it uses all VRAM until it starts offloading and generations of course slow down.
### Steps to Reproduce
Just add two 64 rank lora onto flux schnell fp8
### Debug Logs
```powershell
None
```
### Other
_No response_
Contributor guide
Assessment
This issue has not been assessed yet.