kvcache-ai / kvcache-ai/ktransformers

Pipeline Parallel for ktransformers for multi low-vram GPUs

Open
#1,933 1 comment 0 reactions 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
19.5k
Forks
1.6k
Avg merge
19h 32m
Merged PRs (30d)
27

Description

### Reminder

- [x] I have read the above rules and searched the existing issues.

### Description

Looks like the pipeline parallel for sglang is not well implemented and we are a bit away from sglang main branch, I think we can use well-maintained vllm integrating with sglang parameter to have literal full control over this entire project.

### Pull Request

_No response_

Contributor guide

Open the contributing guide

Research direction

No files, tests, or entry points are named. Start by locating the existing sglang pipeline-parallel implementation and the project’s vLLM integration, then clarify the supported multi-GPU configurations and acceptance criteria for pipeline parallelism on low-VRAM GPUs.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai, distributed-systems
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.