antirez / antirez/ds4

gguf-tools splicer cannot read the MXFP4 GGUF this repo publishes (unsupported GGML tensor type 39)

未关闭 适合新手
#693 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
C
星标
22.4k
派生
2.1k
平均合并
1 天 3 小时
30 天内合并 PR
4

描述

The Flash MXFP4 GGUF (`ds4f-mxfp4`) cannot be used as a donor or base with `gguf-tools/mixed/splice_mixed_expert_layers_gguf.py`, because its `GGML_QUANT_SIZES` table has no entry for type 39.

```
error: unsupported GGML tensor type 39; add it to GGML_QUANT_SIZES
```

Everything else in the splicer already handles it correctly — it deliberately does not require base and donor tensor *types* to match, and it writes each tensor's own `ggml_type` through. Only the block-size table is missing an entry.

One-line fix, using the values from ds4's own table at `ds4.c:2052` (`[39] = {"mxfp4", 32, 17}`):

```python
GGML_QUANT_SIZES = {
...
26: (1, 4, "I32"),
+ 39: (32, 17, "MXFP4"),
}
```

Verified locally with `--dry-run`, splicing MXFP4 routed experts for layers 25-42 onto the IQ2XXS/Q2_K hybrid base:

```
selected donor tensors: 54
selected donor types: MXFP4:54
base tensor payload: 90.88 GiB
mixed tensor payload: 107.76 GiB
```

Motivation, in case it is useful: this produces a quant sized for a single 128 GB Mac. The published Flash quants are 80.8 / 90.9 GiB (fit, leaving ~21 GiB of the 112 GiB wired limit unused) then jump to 145.3 / 153.3 GiB (do not fit). Splicing MXFP4 experts into the upper layers fills that gap at 107.76 GiB.

Happy to open a PR if useful.

贡献指南

打开贡献指南

调研方向

Open gguf-tools/mixed/splice_mixed_expert_layers_gguf.py and locate GGML_QUANT_SIZES; compare the MXFP4 values with ds4.c:2052. Run the reported --dry-run splice for layers 25-42, then confirm type 39 is accepted and the selected donor tensors and payload summary complete without the unsupported-type error.

由索引模型根据 Issue 内容生成。

评估

技术栈
c, python
领域
machine-learning, tooling
Issue 类型
缺陷
难度
1/5
预计耗时
1 小时以内
活跃度
冷清
描述清晰度
描述清楚
新手友好度
92/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。