lmstudio-ai / lmstudio-ai/mlx-engine

Status of MTP for GLM?

Open
#205 4 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
1.2k
Forks
133
Avg merge
21h 6m
Merged PRs (30d)
1

Description

I'm using the 4bit quant of GLM 4.5 Air for MLX. I see in release notes for GLM that it supports Multi-Token-Prediction (MTP), but it sounds different to the speculative decoding that LMStudio has.

I see the LLama cpp doesn't support this yet, but was wondering if LMStudio does, I can't seem to find anything on it.

Basically just wondering if it is already doing it and I just don't realize it, if it is in the works at all, or if it is blocked upstream (Not supported by MLX or something).

Thanks

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no files, tests, or entry points. Begin by determining whether the MLX engine exposes GLM 4.5 Air's Multi-Token-Prediction path or only speculative decoding, then establish whether the work is supported upstream. Done should be a confirmed support status or a clearly scoped implementation path.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.