vllm-project / vllm-project/aibrix

Support Context Cache for Improved Conversation Efficiency

Open
#1,248 2 comments 0 reactions 1 assignee Claimed by @zhengkezhou1 View on GitHub
area/kv-cache help wanted
Dominant language
Go
Stars
5.1k
Forks
697
Avg merge
1d 19h
Merged PRs (30d)
104

Description

### 🚀 Feature Description and Motivation

In many large language model (LLM) scenarios, especially multi-turn conversations or sessions where the user interacts repeatedly with the same context (e.g. chatbots, agents, assistant-like use cases), it’s critical to efficiently re-use past prompt / history information without repeatedly sending the entire conversation back to the model.

Several popular APIs already support explicit context caching or context handles:

- [Anthropic Claude’s prompt caching](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) uses cache identifiers to rehydrate previous contexts.
- [Google Gemini context caching](https://ai.google.dev/gemini-api/docs/caching?lang=python) provides context_cache_id to continue conversations.
- [Moonshot Kimi context caching](https://platform.moonshot.cn/docs/guide/use-context-caching-feature-of-kimi-api) allows explicit reuse of context handles.
- [Volcengine](https://www.volcengine.com/docs/82379/1396491) also offers conversation_id for session reuse.

We’d like to introduce an optional context caching interface in AIBrix, so that:

- Clients can pass in a conversation/session ID or similar handle when making requests.
- AIBrix can reuse already-processed KV cache / embedding context for that session, reducing repeated computation.
- Expose:
- A way to create a new context handle (first request)
- A way to continue using an existing handle (subsequent requests)
- A way to explicitly clear / expire handles (or auto-timeout)

This would likely require:

- Storing partial KV cache (or references) indexed by conversation/session IDs.
- Coordinating with AIBrix’ current GPU memory management and eviction mechanisms.
- Ensuring multi-tenant isolation and clean up on failures.

### Use Case

- New API fields (e.g. context_id, clear_context).
- Internal engine / scheduler support to associate context ID with existing KV cache.
- Metrics to track cache hit/miss rate, and memory usage of stored contexts.

### Proposed Solution

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.