huggingface / huggingface/candle

F8E4M3 CUDA kernels are unreachable: `DType::as_str()` yields `f8e4m3`, the kernels are named `f8_e4m3`

Open
#3,886 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
21k
Forks
1.8k
Avg merge
16h 42m
Merged PRs (30d)
25

Description

## Summary

Every CUDA kernel for `F8E4M3` is declared with an underscore between `f8` and
`e4m3`:

```cuda
// candle-kernels/src/cast.cu:159
CAST_OP_FP8(__nv_fp8_e4m3, float, cast_f8_e4m3_f32)
```

but the Rust side builds kernel names from `DType::as_str()`, which returns the
token *without* that underscore:

```rust
// candle-core/src/dtype.rs:87
Self::F8E4M3 => "f8e4m3",
```

So the lookup asks for `cast_f8e4m3_f32`, the module contains
`cast_f8_e4m3_f32`, and the call fails at runtime with
`CUDA_ERROR_NOT_FOUND` / "named symbol not found".

## Reproduction

Any CUDA device:

```rust
let dev = Device::new_cuda(0)?;
let t = Tensor::zeros((2, 2), DType::F8E4M3, &dev)?;
let _ = t.to_dtype(DType::F32)?; // CUDA_ERROR_NOT_FOUND: named symbol not found
```

## Where the names diverge

| site | name requested | name declared | |
|---|---|---|---|
| `cuda_backend/mod.rs:1578` — `format!("cast_{}_{}", self.dtype().as_str(), dtype.as_str())` | `cast_f8e4m3_f32` | `cast_f8_e4m3_f32` | ✗ |
| `cuda_backend/mod.rs:130` — `kernel_name::(root)` → `{root}_{dtype}` | `affine_f8e4m3`, `ucopy_f8e4m3`, … | `affine_f8_e4m3`, … | ✗ |
| `cuda_backend/mod.rs:2498` — hardcoded `"ucopy_f8e4m3"` | `ucopy_f8e4m3` | `ucopy_f8_e4m3` | ✗ |
| `cuda_backend/mod.rs:1521` — hardcoded `"const_set_f8_e4m3"` | `const_set_f8_e4m3` | `const_set_f8_e4m3` | ✓ |

The only call site that works is the one that spells the name out by hand.
Everything routed through `as_str()` misses.

## Scope

`candle-kernels` declares `F8E4M3` kernels across `affine.cu`, `binary.cu`,
`cast.cu`, `fill.cu`, `indexing.cu`, `unary.cu` and friends — `affine_f8_e4m3`,
`badd/bdiv/bmul/bsub/bmaximum/bminimum_f8_e4m3`, `eq/ge/gt_f8_e4m3`,
`copy2d_f8_e4m3`, `fill_f8_e4m3`, the `gather_*`/`is_*`/`ia_*` family, every
`cast_*` pair. All of them are addressed through the generated name, so none of
them is reachable.

In practice an `F8E4M3` tensor on the CUDA backend cannot be cast, copied,
scaled, compared or indexed. It can be allocated and filled — that path uses
the hardcoded name.

## This is not an architecture problem

Worth stating explicitly, because the failure mode invites the wrong
conclusion. `candle-kernels` does gate `F8E4M3` kernels by architecture, and
the thresholds are not uniform — casts sit under `__CUDA_ARCH__ >= 800` while
`affine_f8_e4m3` and the indexing family sit under `>= 890`. That makes
"symbol not found" look like a compute-capability story, and we spent
considerable time on that reading before checking the name itself.

It is not. `cast_f8_e4m3_f32` is under the `>= 800` guard and is present in the
PTX for `sm_86`, where we hit this. The lookup fails on every architecture,
Blackwell included, because the requested string does not exist in any build.

(The architecture guards remain a real and separate constraint: fixing the
naming makes the `>= 800` kernels reachable everywhere they are compiled, and
leaves the `>= 890` ones legitimately absent below `sm_89`.)

## Diagnosability

Today this surfaces as a bare `DriverError(CUDA_ERROR_NOT_FOUND, "named symbol
not found")` with no indication of which kernel was missing. The name is known
at the failing call site — `get_or_load_func(fn_name, mdl)` — and discarded.
A backtrace does not recover it either, since the error travels up as a value
through `?` and the eventual panic lands wherever the caller unwraps it.

That gap is the subject of a separate report; mentioning it here because
without it this bug is very hard to reach. Naming the kernel in the error turns
this from a multi-day investigation into a one-line diagnosis.

## Fix

Three ways, in increasing order of preference:

1. **Rename the CUDA kernels** `f8_e4m3` → `f8e4m3`. Touches every declaration
across the `.cu` files and requires updating the one hardcoded Rust name
that is currently correct.
2. **Change `as_str()`** to return `"f8_e4m3"`. Note it has other consumers:
`display.rs` prints it as the user-visible dtype of a tensor, and
`metal_backend/mod.rs:1929` builds Metal kernel names from it (Metal
declares no `F8E4M3` kernels today, so only the display output would move).
3. **Separate the two concerns** — give `DType` a dedicated kernel-name token
distinct from its display string, and build CUDA kernel names from that.

Option 3 addresses the actual defect: one string is doing duty both as a
human-facing dtype label and as an ABI-level symbol fragment, and nothing keeps
the two in agreement. Whichever is chosen, the fix is worth pairing with a test
that resolves each declared `F8E4M3` kernel by its generated name, so a future
divergence fails in CI rather than on a user's device.

## Acceptance

- `Tensor::to_dtype` between `F8E4M3` and `F32`/`F16`/`BF16` succeeds on CUDA.
- The generated name for every `F8E4M3` kernel matches a declared symbol, and
something in CI checks that rather than leaving it to runtime.
- `const_set_f8_e4m3` keeps working — whichever direction the rename goes, that
hardcoded name has to move with it.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with candle-core/src/dtype.rs and the CUDA name construction sites in candle-core/src/cuda_backend/mod.rs, then compare those generated names with F8E4M3 declarations across candle-kernels/*.cu. Verify the chosen naming approach preserves display output and const_set_f8_e4m3. Done means CUDA casts succeed, all generated F8E4M3 names match declared symbols, and CI checks the mapping.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
backend, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.