antirez / antirez/ds4

M3 Max 128GB Benchmark Results - ds4-bench on MacBook Pro M3 Max

Open
#917 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
C
Stars
22.3k
Forks
2.1k
Avg merge
1d 3h
Merged PRs (30d)
4

Description

## Hardware

- **Machine**: MacBook Pro, Apple M3 Max, 128 GB unified memory
- **Backend**: Metal (resident model, no SSD streaming)
- **Model**: ds4flash.gguf (q2 quantization, ~91 GiB)
- **ds4 commit**: `c1d4597a` (qa: update DGX Spark host addresses)

## Benchmark Configuration

```bash
./ds4-bench \
-m ds4flash.gguf \
--prompt-file speed-bench/promessi_sposi.txt \
--ctx-start 2048 \
--ctx-max 65536 \
--step-incr 2048 \
--gen-tokens 128
```

## Results

| ctx_tokens | prefill_tps | gen_tps (steady) | kvcache_bytes |
|---:|---:|---:|---:|
| 2048 | 323.28 | 29.35 | 52,184,460 |
| 4096 | 284.69 | 27.04 | 80,373,132 |
| 6144 | 279.15 | 26.55 | 108,561,804 |
| 8192 | 276.19 | 26.50 | 136,750,476 |
| 10240 | 271.39 | 26.42 | 164,939,148 |
| 12288 | 265.27 | 26.58 | 193,127,820 |
| 14336 | 257.17 | 26.39 | 221,316,492 |
| 16384 | 250.37 | 26.22 | 249,505,164 |
| 18432 | 242.65 | 24.52 | 277,693,836 |
| 20480 | 239.42 | 22.87 | 305,882,508 |
| 22528 | 220.16 | 19.94 | 334,071,180 |
| 24576 | 196.02 | 16.53 | 362,259,852 |
| 26624 | 105.19 | 7.75 | 390,448,524 |
| 28672 | 120.96 | 9.22 | 418,637,196 |
| 30720 | 114.46 | 9.43 | 446,825,868 |
| 32768 | 120.13 | 10.41 | 475,014,540 |
| 34816 | 127.20 | 11.42 | 503,203,212 |
| 36864 | 125.72 | 11.70 | 531,391,884 |
| 38912 | 133.03 | 11.85 | 559,580,556 |
| 40960 | 133.38 | 11.94 | 587,769,228 |
| 43008 | 131.64 | 11.87 | 615,957,900 |
| 45056 | 131.29 | 11.87 | 644,146,572 |
| 47104 | 130.34 | 11.88 | 672,335,244 |
| 49152 | 129.83 | 11.99 | 700,523,916 |
| 51200 | 128.31 | 11.88 | 728,712,588 |
| 53248 | 127.53 | 11.96 | 756,901,260 |
| 55296 | 125.76 | 11.37 | 785,089,932 |
| 57344 | 124.33 | 11.85 | 813,278,604 |
| 59392 | 119.48 | 11.85 | 841,467,276 |
| 61440 | 123.07 | 11.77 | 869,655,948 |
| 63488 | 125.52 | 12.28 | 897,844,620 |
| 65536 | 121.52 | 11.00 | 0 |

## Comparison with Official Benchmarks

| Context | This Run (M3 Max 128GB) | Official (M3 Max 128GB) | Delta |
|---|---:|---:|---:|
| Short prompt | **29.35 t/s** | 26.68 t/s | **+10.0%** |
| ~11.7K tokens | **~26.4 t/s** | 21.47 t/s | **+23%** |

## Memory Profile

```
Metal mapped model: 2 overlapping shared buffers
resident model: 90.88 GiB
KV cache: 0.86 GiB (raw 0.36 + compressed 0.50)
buffers: 0.50 GiB
total planned: 92.25 GiB / 128 GB
```

Model is fully resident in Metal memory — no SSD streaming active at these context lengths. Performance drop observed around 26K tokens as KV cache approaches physical memory limits.

## Chart

![M3 Max t/s](https://raw.githubusercontent.com/youssofal/MTPLX/main/docs/ports/deepseek-v4-dspark/BENCHMARKS.md)

CSV and SVG chart saved locally at `speed-bench/your_m3_max.csv` and `speed-bench/your_m3_max_ts.svg`.

## Notes

- Built from commit `c1d4597a` (2026-08-23)
- drift-patch flags: hc_stable=on, norm_unify=on, kv_raw_f32=off
- First token latency: ~38ms at 2K context, ~55ms at 64K
- No SSD streaming; this is the resident model path

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.