lmstudio-ai / lmstudio-ai/mlx-engine
Status of MTP for GLM?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 133
- Avg merge
- 21h 6m
- Merged PRs (30d)
- 1
Description
I'm using the 4bit quant of GLM 4.5 Air for MLX. I see in release notes for GLM that it supports Multi-Token-Prediction (MTP), but it sounds different to the speculative decoding that LMStudio has.
I see the LLama cpp doesn't support this yet, but was wondering if LMStudio does, I can't seem to find anything on it.
Basically just wondering if it is already doing it and I just don't realize it, if it is in the works at all, or if it is blocked upstream (Not supported by MLX or something).
Thanks
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. Begin by determining whether the MLX engine exposes GLM 4.5 Air's Multi-Token-Prediction path or only speculative decoding, then establish whether the work is supported upstream. Done should be a confirmed support status or a clearly scoped implementation path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100