Avoid vLLM copy (doubles the memory)
Open
inference
Performance
vllm
- Dominant language
- Python
- Stars
- 2k
- Forks
- 561
- Avg merge
- 4d 5h
- Merged PRs (30d)
- 145
Description
When vLLM wakes up it uses memory and the cuda IPC also uses memory
Contributor guide
Assessment
This issue has not been assessed yet.