antirez / antirez/h3.c

CUDA backend for Linux/NVIDIA — interest in upstreaming?

オープン
#42 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
C
スター
2.7k
フォーク
210
PR マージ指標
30日以内にマージされた PR はありません

説明

I ported h3.c to Linux ARM64 with a native CUDA backend behind the existing `h3_gpu.h` API (CLI, checkpoint formats and the Metal path unchanged; backend selected at build time).

Measured on an NVIDIA GB10 (DGX Spark, CUDA 13.0), max-quality preset (1024x576, 107 frames, 50 steps/layers):
- full render: 33:37 -> **18:56 wall clock (1.78x)**
- video VAE decode: 1113 s -> 248 s via a tiled F32 attention kernel (4.49x)
- DiT attention: cuDNN SDPA BF16 (~84x on the long shape) with custom tiled fallback
- parity gates: CPU oracles (max_abs <= 2e-5), Compute Sanitizer clean, matched renders SSIM >= 0.999 on the short gate

The branch is ~6.8k lines across 3 commits (portable host/tokenizer, CUDA backend, integration + optimizations). Happy to open a PR if this is of interest, and to split or rework anything.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。