[etLLM] KV Cache as IO by default, as attribute optional
Open
@JacobSzwejbka is already working on this.
Since Feb 13, 2025.
module: llm
triaged
- Dominant language
- Python
- Stars
- 5k
- Forks
- 1.2k
- Avg merge
- 2d 10h
- Merged PRs (30d)
- 581
Description
🚀 The feature, motivation and pitch
For accelerators (HTP and ANE), KV Cache as IO is used to better fit accelerator needs and performance optimization. However, in llama_transformer, KV cache is embedded as module attribute, which is also the running mode for CPU. It brings deviation in the entire flow. For example, we may maintain unnecessarily different runtime logics for different backends.
This issue is to set KV Cache as IO by default, as attribute optional.
- It does not affect CPU performance.
- Exposing KV cache as IO would help user manage KV cache.
- Keep the attribute option for specific use cases where KV cache may not be passed by IO, for example, in Vulkan.
Alternatives
No response
Additional context
No response
RFC (Optional)
No response
cc @mergennachin @cccclai @helunwencser @jackzhxng
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.