vllm-project / vllm-project/aibrix
[Docs]: How to use AiBrix and its KV Management on NPU clusters
- Dominant language
- Go
- Stars
- 5.1k
- Forks
- 697
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 104
Description
### Summary
I will provide documentation and examples for deploying inference services on NPU clusters using AiBrix+ Ascend ecosystem components. At the same time, I will practice the KV cache pooling capability based on the AiBrix and vllm-ascend workflow.
Ascend: https://www.hiascend.com/
how to use: https://gitcode.com/Ascend/mind-cluster/blob/master/docs/zh/scheduling/usage/vllm_best_practice.md
### Motivation
I've noticed that there are some issues in the community asking about how to use AiBrix on Huawei's NPU clusters. Now, our team hopes to offer some assistance to enable AiBrix community users to have a better experience when working on NPU clusters.
#1861 #1753 #1476 #1858
### Proposed Change
We have already practiced using Aibrix on the NPU cluster without requiring any code adaptation, which thanks to the well-structured architecture of the Aibrix framework. However, we have not yet verified the KVCache capability, which may involve some code adaptation.
### Alternatives Considered
_No response_
Contributor guide
Research direction
Start with the linked Ascend vllm-ascend best-practice guide and review issues #1861, #1753, #1476, and #1858 for the questions the documentation should address. Document deployment examples for AiBrix inference services on NPU clusters and the KV cache pooling workflow; done means users can follow the examples and the KV cache capability is verified or its limitation is recorded.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai-infra-agents, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 45/100