antirez / antirez/ds4

GLM 5.2 - Distributed Inference Fail

Đang mở
#505 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

### Summary

Distributed inference across two Apple Silicon machines fails during prompt
processing with:

ds4: prompt processing failed: metal GLM layer-slice evaluation failed at pos 0

The failure happens immediately, at position 0, i.e. before any token is
generated (prefill of the very first chunk).

### Environment

- Backend: Metal
- Distributed inference: 2 machines
- Coordinator: Mac Studio M3 Ultra, 256 GB RAM
- Worker: Mac Studio M3 Ultra, 96 GB RAM
- macOS version: 27 Beta 2
- Model: GLM 5.2, quant: Q2_K

### Steps to reproduce

1. Start the worker on the 96 GB machine.
2. Start the coordinator on the 256 GB machine with the distributed route above.
3. Send any prompt (the error triggers at pos 0, so even a trivial prompt reproduces it).

### Expected behavior

Prefill of the first chunk completes and generation starts.

### Actual behavior

Prompt processing aborts at position 0:

ds4: prompt processing failed: metal GLM layer-slice evaluation failed at pos 0

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.