InternLM / InternLM/lmdeploy

【Design Questinon】any plan to decouple batching and cache from llama?

Open
#476 1 comment 0 reactions 1 assignee Claimed by @lzhangzz View on GitHub
backlog
Dominant language
Python
Stars
8.1k
Forks
748
Avg merge
6d 2h
Merged PRs (30d)
54

Description

Is there any reason why the batching and cache manager are implemented inside llama? Those looks generic functionalities, and better not mixed with the llama code. It looks misleading that turbomind only supports llama.

And is there any plan the abstract and decouple those functionalities from llama?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.