Cross-Vendor Multi-GPU Support via Vulkan Backend
Open
User Support
- Dominant language
- Python
- Stars
- 133k
- Forks
- 15.7k
- Avg merge
- 1d 6h
- Merged PRs (30d)
- 155
Description
I realise that this is a far stretch, but a vulkan backend which would allow for multi-gpu inference would be awesome. With new models like Flux, inference requires more than 12gb at fp8. Vulkan would solve cross nvidia-AMD setups and had been implemented in llama.cpp (would be different, but its still torch). I got llama inference working on my 3060 12gb rx 6750xt setup with vulkan, so it would be great if something like this was possible.
While similar requests have been brought up in the past, it was unclear to me whether or not they were meant for multi-gpu inference for singular images in the interest of vram.
Contributor guide
Assessment
This issue has not been assessed yet.