ml-explore / ml-explore/mlx-examples
Higher Speed High-Context processing w/ Minference
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 9k
- Forks
- 1.2k
- PR merge metrics
- No merged PRs in 30d
Description
Would it be possible to implement the KV architecture from MInference to speed up long-context inputs (ex. for Qwen2.5-1M)?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no target file, test, or entry point. Read the linked MInference project and inspect the Python MLX examples to identify the long-context processing path; done should include the requested KV architecture and faster processing for inputs such as Qwen2.5-1M.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100