antirez / antirez/ds4

tensor parallel with uneven number of GPU's?

Đang mở
#624 2 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

> --cuda-tensor-parallel splits DeepSeek V4 Flash tensor and routed-expert work across an even number of GPUs.

Is this "even number of GPUs" limitation actually present in the ds4 code? Does it need to be? I know that vllm has such a limitation, but I don't believe llama-server does, for example.

It would be great to get tensor-parallel on (my) 5x3090 or maybe (someone else's) 3x32GB, for example :)

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.