kvcache-ai / kvcache-ai/ktransformers
Pipeline Parallel for ktransformers for multi low-vram GPUs
- Dominant language
- Python
- Stars
- 19.5k
- Forks
- 1.6k
- Avg merge
- 19h 32m
- Merged PRs (30d)
- 27
Description
### Reminder
- [x] I have read the above rules and searched the existing issues.
### Description
Looks like the pipeline parallel for sglang is not well implemented and we are a bit away from sglang main branch, I think we can use well-maintained vllm integrating with sglang parameter to have literal full control over this entire project.
### Pull Request
_No response_
Contributor guide
Research direction
No files, tests, or entry points are named. Start by locating the existing sglang pipeline-parallel implementation and the project’s vLLM integration, then clarify the supported multi-GPU configurations and acceptance criteria for pipeline parallelism on low-VRAM GPUs.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, distributed-systems
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100