lmstudio-ai / lmstudio-ai/mlx-engine

Qwen3 cache wastage

Open
#176 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.2k
Forks
133
Avg merge
21h 6m
Merged PRs (30d)
1

Description

Qwen3 has the following prompt format (shown with thinking disabled):

```
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What are the prime numbers between 1 and 100?<|im_end|>
<|im_start|>assistant

```

The model receives this as input and generates its response, resulting in e.g.

```
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What are the prime numbers between 1 and 100?<|im_end|>
<|im_start|>assistant

I have no idea.<|im_end|>
```

When the user submits another message, the `` block moves to the end of context again, which triggers the model to throw away its entire previous generation from the KV cache since it is no longer part of a common prefix with the previous generation:

```
<|im_start|>system
You are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>
<|im_start|>user
What are the prime numbers between 1 and 100?<|im_end|>
<|im_start|>assistant
I have no idea.<|im_end|>
<|im_start|>user
Aren't you supposed to know these things?<|im_end|>
<|im_start|>assistant

```

Here, the block

```

I have no idea.<|im_end|>
```

is evicted from the cache and the line `I have no idea.<|im_end|>` must be reprocessed. This is not much of a problem in this dummy example, but with a very long assistant response, this can result in thousands of tokens being unnecessarily reprocessed. This effectively causes every assistant-generated token to be processed twice, once the first time it's generated and once afterward, which adds up to be very costly over long sequences with big Qwen3 models.

Contributor guide

Open the contributing guide

Research direction

The issue does not name files or tests; start by locating Qwen3 prompt construction and KV-cache prefix handling in the engine. Reproduce a multi-turn Qwen3 conversation with thinking disabled, then verify that the prior assistant response remains reusable in the cache instead of being reprocessed.

Written by the indexing model from the issue text.

Assessment

Domain
ai-infra-agents, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.