antirez / antirez/ds4

M3 Max 128GB Benchmark Results - ds4-bench on MacBook Pro M3 Max

Abierto
#917 4 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
C
Estrellas
22.3k
Forks
2.1k
Merge medio
1 d 3 h
PR fusionados (30 d)
4

Descripción

## Hardware

- **Machine**: MacBook Pro, Apple M3 Max, 128 GB unified memory
- **Backend**: Metal (resident model, no SSD streaming)
- **Model**: ds4flash.gguf (q2 quantization, ~91 GiB)
- **ds4 commit**: `c1d4597a` (qa: update DGX Spark host addresses)

## Benchmark Configuration

```bash
./ds4-bench \
-m ds4flash.gguf \
--prompt-file speed-bench/promessi_sposi.txt \
--ctx-start 2048 \
--ctx-max 65536 \
--step-incr 2048 \
--gen-tokens 128
```

## Results

| ctx_tokens | prefill_tps | gen_tps (steady) | kvcache_bytes |
|---:|---:|---:|---:|
| 2048 | 323.28 | 29.35 | 52,184,460 |
| 4096 | 284.69 | 27.04 | 80,373,132 |
| 6144 | 279.15 | 26.55 | 108,561,804 |
| 8192 | 276.19 | 26.50 | 136,750,476 |
| 10240 | 271.39 | 26.42 | 164,939,148 |
| 12288 | 265.27 | 26.58 | 193,127,820 |
| 14336 | 257.17 | 26.39 | 221,316,492 |
| 16384 | 250.37 | 26.22 | 249,505,164 |
| 18432 | 242.65 | 24.52 | 277,693,836 |
| 20480 | 239.42 | 22.87 | 305,882,508 |
| 22528 | 220.16 | 19.94 | 334,071,180 |
| 24576 | 196.02 | 16.53 | 362,259,852 |
| 26624 | 105.19 | 7.75 | 390,448,524 |
| 28672 | 120.96 | 9.22 | 418,637,196 |
| 30720 | 114.46 | 9.43 | 446,825,868 |
| 32768 | 120.13 | 10.41 | 475,014,540 |
| 34816 | 127.20 | 11.42 | 503,203,212 |
| 36864 | 125.72 | 11.70 | 531,391,884 |
| 38912 | 133.03 | 11.85 | 559,580,556 |
| 40960 | 133.38 | 11.94 | 587,769,228 |
| 43008 | 131.64 | 11.87 | 615,957,900 |
| 45056 | 131.29 | 11.87 | 644,146,572 |
| 47104 | 130.34 | 11.88 | 672,335,244 |
| 49152 | 129.83 | 11.99 | 700,523,916 |
| 51200 | 128.31 | 11.88 | 728,712,588 |
| 53248 | 127.53 | 11.96 | 756,901,260 |
| 55296 | 125.76 | 11.37 | 785,089,932 |
| 57344 | 124.33 | 11.85 | 813,278,604 |
| 59392 | 119.48 | 11.85 | 841,467,276 |
| 61440 | 123.07 | 11.77 | 869,655,948 |
| 63488 | 125.52 | 12.28 | 897,844,620 |
| 65536 | 121.52 | 11.00 | 0 |

## Comparison with Official Benchmarks

| Context | This Run (M3 Max 128GB) | Official (M3 Max 128GB) | Delta |
|---|---:|---:|---:|
| Short prompt | **29.35 t/s** | 26.68 t/s | **+10.0%** |
| ~11.7K tokens | **~26.4 t/s** | 21.47 t/s | **+23%** |

## Memory Profile

```
Metal mapped model: 2 overlapping shared buffers
resident model: 90.88 GiB
KV cache: 0.86 GiB (raw 0.36 + compressed 0.50)
buffers: 0.50 GiB
total planned: 92.25 GiB / 128 GB
```

Model is fully resident in Metal memory — no SSD streaming active at these context lengths. Performance drop observed around 26K tokens as KV cache approaches physical memory limits.

## Chart

![M3 Max t/s](https://raw.githubusercontent.com/youssofal/MTPLX/main/docs/ports/deepseek-v4-dspark/BENCHMARKS.md)

CSV and SVG chart saved locally at `speed-bench/your_m3_max.csv` and `speed-bench/your_m3_max_ts.svg`.

## Notes

- Built from commit `c1d4597a` (2026-08-23)
- drift-patch flags: hc_stable=on, norm_unify=on, kv_raw_f32=off
- First token latency: ~38ms at 2K context, ~55ms at 64K
- No SSD streaming; this is the resident model path

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.