lmstudio-ai / lmstudio-ai/mlx-engine
Qwen3 cache wastage
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 133
- Avg merge
- 21h 6m
- Merged PRs (30d)
- 1
Description
Qwen3 has the following prompt format (shown with thinking disabled):
```
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What are the prime numbers between 1 and 100?<|im_end|>
<|im_start|>assistant
```
The model receives this as input and generates its response, resulting in e.g.
```
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What are the prime numbers between 1 and 100?<|im_end|>
<|im_start|>assistant
I have no idea.<|im_end|>
```
When the user submits another message, the `` block moves to the end of context again, which triggers the model to throw away its entire previous generation from the KV cache since it is no longer part of a common prefix with the previous generation:
```
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What are the prime numbers between 1 and 100?<|im_end|>
<|im_start|>assistant
I have no idea.<|im_end|>
<|im_start|>user
Aren't you supposed to know these things?<|im_end|>
<|im_start|>assistant
```
Here, the block
```
I have no idea.<|im_end|>
```
is evicted from the cache and the line `I have no idea.<|im_end|>` must be reprocessed. This is not much of a problem in this dummy example, but with a very long assistant response, this can result in thousands of tokens being unnecessarily reprocessed. This effectively causes every assistant-generated token to be processed twice, once the first time it's generated and once afterward, which adds up to be very costly over long sequences with big Qwen3 models.
Contributor guide
Research direction
The issue does not name files or tests; start by locating Qwen3 prompt construction and KV-cache prefix handling in the engine. Reproduce a multi-turn Qwen3 conversation with thinking disabled, then verify that the prior assistant response remains reusable in the cache instead of being reprocessed.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai-infra-agents, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100