antirez / antirez/h3.c

M1 Ultra profile + two negative results (TensorOps unavailable, M5 gates are no-ops)

Aperta
#17 1 commento 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
C
Stelle
2.7k
Fork
210
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

First, thanks for open-sourcing this — it builds and runs cleanly on M1 Ultra with zero warnings, and it is roughly **2.15× faster than the Swift/MLX path** on the same machine while using **44% less memory**, at full BF16 rather than 4-bit. That is a big deal for people on older Apple Silicon.

I did not find M1/M2-series numbers anywhere in the issues (#5 and #8 cover M4 Max, #15 mentions M3 Ultra), so here is a complete profile from an M1 Ultra, plus two negative results that may save others some time.

## Environment

- Apple M1 Ultra, 128 GB, macOS 26.2
- `h3.c` at `8974cc0`
- Official `MiniMaxAI/MiniMax-H3` snapshot, `FL2VA/` only

```
Device: Apple M1 Ultra (applegpu_g13d)
physical memory 128.0 GiB
recommended GPU set 107.5 GiB
max Metal buffer 80.6 GiB
Apple GPU family 7
Metal 4 yes
unified memory yes
```

## Profile — 512×512, 22 frames, 20 steps, `--layers 45 --reuse 2`

Same preset as the README's balanced example (~16.7 s denoise on M5 Max):

| Stage | wall | encode | wait | peak |
|---|---|---|---|---|
| Qwen text encoder | 10.5 s | 0.003 s | 5.4 s | 2.7 GiB |
| H3 DiT load | 18.3 s | 1.6 s | 1.7 s | 32.8 GiB |
| **H3 DiT Euler denoise** | **86.6 s** | 56.7 s | 29.8 s | 32.8 GiB |
| audio VAE decode | 2.6 s | 2.3 s | 0.2 s | 0.3 GiB |
| video VAE decode | 7.3 s | 0.9 s | 4.3 s | 9.5 GiB |
| **total** | **~125 s** | | | **32.8 GiB** |

So denoise is **≈5.2× slower than M5 Max** on the identical preset. `encode` dominates over `wait`, i.e. compute-bound rather than bandwidth-bound — which matches expectations, since M1 Ultra actually has *higher* memory bandwidth (800 GB/s) than M5 Max but no per-core matrix engine.

## Profile — 768×416, `--steps 15 --layers 45 --reuse 2`

Two frame counts, showing denoise scales close to linearly:

| frames | duration | denoise | total wall | peak |
|---|---|---|---|---|
| 141 | 5.875 s | 822.9 s | **15 min 14 s** | 34.1 GiB |
| 243 | 10.125 s | 1930.1 s | **34 min 28 s** | 35.2 GiB |

One observation: at 243 frames the video VAE decode was **105.3 s wall with only 0.9 s encode** — 101.7 s of `wait`. Decode appears almost entirely wait-bound at higher frame counts on this hardware.

## Negative result 1 — TensorOps is a hard hardware limit, not a gate

I wanted to check whether the `is_m5()` name match in `h3_gpu.m` was leaving performance on the table, so I compiled a minimal probe directly:

```objc
NSString *src = @"#include \n#include \n"
"using namespace metal;\nusing namespace mpp::tensor_ops;\n"
"kernel void probe(device float* o [[buffer(0)]]) { o[0]=1.0f; }\n";
id lib = [dev newLibraryWithSource:src options:o error:&err];
```

Result on M1 Ultra:

```
Device: Apple M1 Ultra
Metal4 family: YES
TensorOps compile: FAILED
Reason: program_source:4:17: error: use of undeclared identifier 'mpp'
```

So `supportsFamily:MTLGPUFamilyMetal4` returns YES while `mpp::tensor_ops` does not exist at all — the framework is available but the namespace is gated on the GPU neural accelerators introduced in M5/A19. Forcing `is_m5()` true would just hit the existing fallback at `h3_gpu.m:381-393`. **Your name-based check is semantically correct**; there is no free performance hiding behind it.

## Negative result 2 — the three M5 software gates are no-ops here

Three gates select M5 behaviour by device name and are overridable by env var:

- `H3_GPU_SAMPLER` (`h3_dit.c:2584`) — GPU sampler, M5-only by default
- `H3_QWEN_PREFETCH_DEPTH` (`h3_text_encoder.c:137`) — 3 on M5, 2 otherwise
- `H3_DIT_COMMAND_BLOCKS` (`h3_dit.c:216`) — `blocks*3/5` on M5, 30 otherwise

A/B on 512×512 / 22 frames / 20 steps, DiT total wall:

| run | config | DiT total |
|---|---|---|
| A | defaults (cold) | 106.88 s |
| B | `H3_GPU_SAMPLER=1` | 103.49 s |
| C | `H3_QWEN_PREFETCH_DEPTH=3` | 103.46 s |
| D | B + C | 103.55 s |
| E | D + `H3_DIT_COMMAND_BLOCKS=27` | 103.39 s |

B–E span 0.16 s — noise. The 3.4 s gap to A is cold-start (first-run shader compile), not the flags: run C did **not** enable the GPU sampler yet matched B exactly. `H3_DIT_COMMAND_BLOCKS=27` did change submissions (66 → 77) with no timing effect. Prefetch depth 3 only raised text-encoder peak from 2.7 to 3.6 GiB with no speedup.

## The actual question

Given that the compute path is closed on pre-M5 hardware, is there anything on the roadmap targeting the parts that *aren't* matrix-engine dependent? Specifically:

1. **The `wait` fraction.** Denoise showed ~34% wait (29.8 s of 86.6 s), and 243-frame video decode was ~97% wait. If any of that is schedulable rather than inherent, it would be the largest available win on this class of machine.
2. **Whether `--layers`/`--reuse` defaults should differ pre-M5.** The fast preset is clearly the practical operating point here; the exact preset (`--layers 50 --reuse 1 --steps 20`) came out at ~4.5 min/step at 243 frames, which is not usable.

Entirely understood if pre-M5 is out of scope — the README is explicit that the current work is M3 Max and M5 Max. In that case please treat this as a data point for anyone else landing here with an Ultra, and feel free to close.

Happy to run further profiles or test patches on this hardware if useful.

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.