GLM 5.2 - Distributed Inference Fail
- 主要语言
- C
- 星标
- 22.3k
- 派生
- 2.1k
- 平均合并
- 1 天 3 小时
- 30 天内合并 PR
- 4
描述
### Summary
Distributed inference across two Apple Silicon machines fails during prompt
processing with:
ds4: prompt processing failed: metal GLM layer-slice evaluation failed at pos 0
The failure happens immediately, at position 0, i.e. before any token is
generated (prefill of the very first chunk).
### Environment
- Backend: Metal
- Distributed inference: 2 machines
- Coordinator: Mac Studio M3 Ultra, 256 GB RAM
- Worker: Mac Studio M3 Ultra, 96 GB RAM
- macOS version: 27 Beta 2
- Model: GLM 5.2, quant: Q2_K
### Steps to reproduce
1. Start the worker on the 96 GB machine.
2. Start the coordinator on the 256 GB machine with the distributed route above.
3. Send any prompt (the error triggers at pos 0, so even a trivial prompt reproduces it).
### Expected behavior
Prefill of the first chunk completes and generation starts.
### Actual behavior
Prompt processing aborts at position 0:
ds4: prompt processing failed: metal GLM layer-slice evaluation failed at pos 0
贡献指南
评估
这个 Issue 还没有评估数据。