pytorch / pytorch/executorch

[etLLM] KV Cache as IO by default, as attribute optional

Open
#8,417 1 comment 0 reactions 1 assignee View on GitHub

@JacobSzwejbka is already working on this.

Since Feb 13, 2025.

module: llm triaged
Dominant language
Python
Stars
5k
Forks
1.2k
Avg merge
2d 10h
Merged PRs (30d)
581

Description

🚀 The feature, motivation and pitch

For accelerators (HTP and ANE), KV Cache as IO is used to better fit accelerator needs and performance optimization. However, in llama_transformer, KV cache is embedded as module attribute, which is also the running mode for CPU. It brings deviation in the entire flow. For example, we may maintain unnecessarily different runtime logics for different backends.
This issue is to set KV Cache as IO by default, as attribute optional.

  • It does not affect CPU performance.
  • Exposing KV cache as IO would help user manage KV cache.
  • Keep the attribute option for specific use cases where KV cache may not be passed by IO, for example, in Vulkan.
Alternatives

No response

Additional context

No response

RFC (Optional)

No response

cc @mergennachin @cccclai @helunwencser @jackzhxng

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.