michaelfeil / michaelfeil/infinity
Tensor-parallelism for multi-gpu support
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 2.9k
- Forks
- 206
- PR merge metrics
- No merged PRs in 30d
Description
### Feature request
Being able to split models into multiple GPUs, as with vllm/aphrodite engine for LLMs.
### Motivation
It would be extremely helpful to be able to split larger models into multiple GPUs.
Also, without TP, one GPU loses lots of vram and the other does not, making it impossible to use tensor parallelism on another program at the same time. (Without losing as much VRAM on the non-utilized GPU)
### Your contribution
communicating the feature
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. Start by locating the existing model-serving and GPU allocation paths, then compare the requested behavior with tensor parallelism in vllm or aphrodite; done means larger models can be split across multiple GPUs while balancing VRAM use.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100