[Feature]: Conditional Disaggregation Based on the KV Availability and/or ISL limits
@laikhtewari is already working on this.
Since Dec 30, 2025.
- Dominant language
- Python
- Stars
- 14.7k
- Forks
- 2.8k
- Avg merge
- 2d 23h
- Merged PRs (30d)
- 489
Description
🚀 The feature, motivation and pitch
When serving a model, that benefits from disaggregation, it often also benefits from the KV-aware routing features, offered by Dynamo and other inference frameworks.
However, there's more efficiency to gain, when both of these features are enabled if the disaggregation is happening conditionally.
Let's take as an example the usual decoding-first setup. The new request is coming with 10000 input tokens, but 9950 of them already have KV$ computed and stored in the system (50 remaining to prefill)
Currently, even if there is a KV-hit, the disaggregation is still happening and the 9950 KV$ has to be transferred first from the CPU memory to the prefill worker, and then from the prefill worker to the decoding worker.
If the disaggregation could be done conditionally, by the number of tokens that really have to be prefilled, considering the available KV$, the decoding node might have initiated the load of the 9950 KV$ from the CPU memory to itself, and then running the prefill for 50 tokens, which shouldn't have any dramatic impact on the decoding node performance.
cc @ltalal
Alternatives
No response
Additional context
No response
Before submitting a new issue...
- Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.