antirez / antirez/ds4

M3 Max 128GB Benchmark Results - ds4-bench on MacBook Pro M3 Max

Đang mở
#917 4 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
C
Star
22.3k
Fork
2.1k
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
4

Mô tả

## Hardware

- **Machine**: MacBook Pro, Apple M3 Max, 128 GB unified memory
- **Backend**: Metal (resident model, no SSD streaming)
- **Model**: ds4flash.gguf (q2 quantization, ~91 GiB)
- **ds4 commit**: `c1d4597a` (qa: update DGX Spark host addresses)

## Benchmark Configuration

```bash
./ds4-bench \
-m ds4flash.gguf \
--prompt-file speed-bench/promessi_sposi.txt \
--ctx-start 2048 \
--ctx-max 65536 \
--step-incr 2048 \
--gen-tokens 128
```

## Results

| ctx_tokens | prefill_tps | gen_tps (steady) | kvcache_bytes |
|---:|---:|---:|---:|
| 2048 | 323.28 | 29.35 | 52,184,460 |
| 4096 | 284.69 | 27.04 | 80,373,132 |
| 6144 | 279.15 | 26.55 | 108,561,804 |
| 8192 | 276.19 | 26.50 | 136,750,476 |
| 10240 | 271.39 | 26.42 | 164,939,148 |
| 12288 | 265.27 | 26.58 | 193,127,820 |
| 14336 | 257.17 | 26.39 | 221,316,492 |
| 16384 | 250.37 | 26.22 | 249,505,164 |
| 18432 | 242.65 | 24.52 | 277,693,836 |
| 20480 | 239.42 | 22.87 | 305,882,508 |
| 22528 | 220.16 | 19.94 | 334,071,180 |
| 24576 | 196.02 | 16.53 | 362,259,852 |
| 26624 | 105.19 | 7.75 | 390,448,524 |
| 28672 | 120.96 | 9.22 | 418,637,196 |
| 30720 | 114.46 | 9.43 | 446,825,868 |
| 32768 | 120.13 | 10.41 | 475,014,540 |
| 34816 | 127.20 | 11.42 | 503,203,212 |
| 36864 | 125.72 | 11.70 | 531,391,884 |
| 38912 | 133.03 | 11.85 | 559,580,556 |
| 40960 | 133.38 | 11.94 | 587,769,228 |
| 43008 | 131.64 | 11.87 | 615,957,900 |
| 45056 | 131.29 | 11.87 | 644,146,572 |
| 47104 | 130.34 | 11.88 | 672,335,244 |
| 49152 | 129.83 | 11.99 | 700,523,916 |
| 51200 | 128.31 | 11.88 | 728,712,588 |
| 53248 | 127.53 | 11.96 | 756,901,260 |
| 55296 | 125.76 | 11.37 | 785,089,932 |
| 57344 | 124.33 | 11.85 | 813,278,604 |
| 59392 | 119.48 | 11.85 | 841,467,276 |
| 61440 | 123.07 | 11.77 | 869,655,948 |
| 63488 | 125.52 | 12.28 | 897,844,620 |
| 65536 | 121.52 | 11.00 | 0 |

## Comparison with Official Benchmarks

| Context | This Run (M3 Max 128GB) | Official (M3 Max 128GB) | Delta |
|---|---:|---:|---:|
| Short prompt | **29.35 t/s** | 26.68 t/s | **+10.0%** |
| ~11.7K tokens | **~26.4 t/s** | 21.47 t/s | **+23%** |

## Memory Profile

```
Metal mapped model: 2 overlapping shared buffers
resident model: 90.88 GiB
KV cache: 0.86 GiB (raw 0.36 + compressed 0.50)
buffers: 0.50 GiB
total planned: 92.25 GiB / 128 GB
```

Model is fully resident in Metal memory — no SSD streaming active at these context lengths. Performance drop observed around 26K tokens as KV cache approaches physical memory limits.

## Chart

![M3 Max t/s](https://raw.githubusercontent.com/youssofal/MTPLX/main/docs/ports/deepseek-v4-dspark/BENCHMARKS.md)

CSV and SVG chart saved locally at `speed-bench/your_m3_max.csv` and `speed-bench/your_m3_max_ts.svg`.

## Notes

- Built from commit `c1d4597a` (2026-08-23)
- drift-patch flags: hc_stable=on, norm_unify=on, kv_raw_f32=off
- First token latency: ~38ms at 2K context, ~55ms at 64K
- No SSD streaming; this is the resident model path

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.