antirez / antirez/ds4

tensor parallel with uneven number of GPU's?

Aperta
#624 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
C
Stelle
22.3k
Fork
2.1k
Merge medio
1g 3h
PR unite (30g)
4

Descrizione

> --cuda-tensor-parallel splits DeepSeek V4 Flash tensor and routed-expert work across an even number of GPUs.

Is this "even number of GPUs" limitation actually present in the ds4 code? Does it need to be? I know that vllm has such a limitation, but I don't believe llama-server does, for example.

It would be great to get tensor-parallel on (my) 5x3090 or maybe (someone else's) 3x32GB, for example :)

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.