Comfy-Org / Comfy-Org/comfy-aimdo

HIP/ROCm support for AMD GPUs — working port on RX 9070 (gfx1201)

Open
#29 13 comments 1 reaction 0 assignees View on GitHub
Dominant language
C
Stars
67
Forks
39
Avg merge
1d 25m
Merged PRs (30d)
10

Description

Hi there! 👋

I stumbled across the [Dynamic VRAM blog post](https://blog.comfy.org/p/dynamic-vram-in-comfyui-saving-local) today and got really excited — this is exactly what I've been needing for my AMD setup. The post mentions AMD support is planned for the future, but I couldn't wait, so I fired up my slop generator (Claude Code) and about an hour later it came back with a working HIP port.

## What I did

The entire port lives in a single file change to `src/plat.h` — an `#ifdef __HIP_PLATFORM_AMD__` block at the top that maps all CUDA Driver API types and functions to their HIP equivalents. All `.c` files remain completely unmodified.

**Key points:**
- `CUdeviceptr` kept as `unsigned long long` (matching CUDA) with cast macros for HIP APIs that expect `void*` — preserves all upstream pointer arithmetic
- `cuMemAllocHost` wrapped to supply HIP's extra `flags` parameter
- `cuGetErrorString` wrapped to adapt HIP's direct-return API style
- `pyt-cu-plug-alloc-async.c` excluded from the build (the `cudaMalloc`/`cudaFree` symbol interposition is not portable to HIP, but the VBAR path that ComfyUI uses doesn't depend on it)

## Test results (RX 9070, gfx1201, ROCm 7.2)

All HIP VMM APIs work correctly:
- `hipMemAddressReserve` — zero VRAM cost, as expected
- `hipMemCreate` / `hipMemMap` / `hipMemSetAccess` — physical VRAM allocated and accessible
- `hipMemUnmap` / `hipMemRelease` — clean deallocation
- Full VBAR allocate → fault → unpin lifecycle works

ComfyUI startup log:
```
comfy-aimdo inited for GPU: AMD Radeon RX 9070 (VRAM: 16304 MB)
DynamicVRAM support detected and enabled
```

Workflows that previously crashed with OOM (multi-model, full 1 megapixel resolution) now complete successfully with dynamic weight eviction working as intended.

## The fork

I've pushed the change to a branch on my fork: [spheenik/comfy-aimdo@hip-rocm-support](https://github.com/spheenik/comfy-aimdo/compare/master...spheenik:comfy-aimdo:hip-rocm-support)

Happy to open a PR if you're interested. I understand you might want to approach this differently for an official implementation (CI, wheel builds, etc.) — just wanted to share that the HIP VMM APIs are there and working on RDNA 4, in case it's useful for your roadmap.

## Build instructions (for anyone who wants to try)

```bash
gcc -shared -fPIC -O2 -D__HIP_PLATFORM_AMD__ -Isrc \
-I/opt/rocm/include -L/opt/rocm/lib \
src/vrambuf.c src/model-vbar.c src/control.c src/hostbuf.c \
src/debug.c src/pyt-cu-plug-alloc.c src-posix/model-mmap.c \
-lamdhip64 -o aimdo.so
```

Then replace the CUDA `aimdo.so` inside the `comfy_aimdo` pip package directory and start ComfyUI with `--enable-dynamic-vram`.

Thanks for the great work on AIMDO — it's a game changer!

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.