antirez / antirez/ds4

GLM 5.2 - Distributed Inference Fail

未关闭
#505 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
C
星标
22.3k
派生
2.1k
平均合并
1 天 3 小时
30 天内合并 PR
4

描述

### Summary

Distributed inference across two Apple Silicon machines fails during prompt
processing with:

ds4: prompt processing failed: metal GLM layer-slice evaluation failed at pos 0

The failure happens immediately, at position 0, i.e. before any token is
generated (prefill of the very first chunk).

### Environment

- Backend: Metal
- Distributed inference: 2 machines
- Coordinator: Mac Studio M3 Ultra, 256 GB RAM
- Worker: Mac Studio M3 Ultra, 96 GB RAM
- macOS version: 27 Beta 2
- Model: GLM 5.2, quant: Q2_K

### Steps to reproduce

1. Start the worker on the 96 GB machine.
2. Start the coordinator on the 256 GB machine with the distributed route above.
3. Send any prompt (the error triggers at pos 0, so even a trivial prompt reproduces it).

### Expected behavior

Prefill of the first chunk completes and generation starts.

### Actual behavior

Prompt processing aborts at position 0:

ds4: prompt processing failed: metal GLM layer-slice evaluation failed at pos 0

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。