How to choose between models for the 512GB max m3ultra
Aperta
- Lingua principale
- C
- Stelle
- 22.4k
- Fork
- 2.1k
- Merge medio
- 1g 3h
- PR unite (30g)
- 4
Descrizione
I can see there is different prefill and tokens/sec but how much smarter is pro-imatrix vs q4-imatrix? also how about q2-q4-imatrix? any insight is much appreciated.
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.