deepseek-ai / deepseek-ai/DeepGEMM

[Bug]: NVRTC JIT compilation fails on CUDA 12.8 (smxx_clean_logits.cuh: expression must have a constant value)

Open Beginner friendly
#295 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Cuda
Stars
7.8k
Forks
1.3k
Avg merge
3d 7h
Merged PRs (30d)
3

Description

### Describe the bug
When running DeepGEMM on CUDA 12.8, the `fp8_mqa_logits` kernel fails to JIT-compile via NVRTC. The compiler throws an error stating that `cute::numeric_limits::infinity()` cannot be used as a `constexpr`.

### Error Log
```text
NVRTC log: "kernel.cu": creating precompiled header file "kernel.pch"
/usr/local/lib/python3.12/dist-packages/deep_gemm/include/deep_gemm/impls/smxx_clean_logits.cuh(17): error: expression must have a constant value
constexpr float neg_inf = -cute::numeric_limits::infinity();
^
/usr/local/cuda/include/cuda/std/detail/libcxx/include/limits(686): note #2703-D: cannot call non-constexpr function "cuda::std::__4::__libcpp_numeric_limits::infinity"
```

### Root Cause & Investigation
In CUDA 12.8 device code, the underlying implementation of `cuda::std::numeric_limits::infinity()` inside `libcxx/include/limits` is missing the `constexpr` qualifier under NVRTC's stripped-down compilation environment.

### What we tried before finding the fix
- Cleared JIT cache (`/root/.deep_gemm`, `/workspace/.deep_gemm_cache/cache`).
- Reinstalled DeepGEMM non-editable (`pip install --no-build-isolation .`).
- Removed stale editable install artifacts (`.egg-info`, `.pth` files).
- Set `CPLUS_INCLUDE_PATH` to CUTLASS headers (NVRTC ignores it).
- Replaced with standard built-ins (`__builtin_inff()`) — NVRTC strips host-side built-ins and throws an undefined identifier error.
- Replaced with native CUDA intrinsics (`__int_as_float(0xff800000)`) — NVRTC refuses to evaluate it as a constant expression.
- Implemented a PyTorch `try/except` fallback using `torch.einsum` to bypass the crashed kernel — this successfully unblocked the model run, confirming the JIT failure was the only roadblock.
- **Only a `-1e38f` raw literal worked as a true compile-time constant for NVRTC.**

### Proposed Fix / Workaround
To maintain compatibility with CUDA 12.8 without requiring users to upgrade to 12.9, replacing the standard library call with a raw float literal fixes the NVRTC compilation instantly.

In `deep_gemm/include/deep_gemm/impls/smxx_clean_logits.cuh` (Line 17):
**Change from:**
`constexpr float neg_inf = -cute::numeric_limits::infinity();`

**Change to:**
`constexpr float neg_inf = -1e38f;`

Since `-1e38f` mathematically acts as negative infinity to mask out attention scores during the Softmax step, it preserves accuracy while bypassing the NVRTC `constexpr` limitation.

### Environment
* **DeepGEMM Version:** Version 2.3.0, commit d30fc36
* **GPU:** 1x NVIDIA H100 SXM (80GB HBM3, SM90)
* **System Resources:** 26 vCPUs, 251 GB Memory
* **Environment:** Docker container (Template: `runpod-torch-v280`)
* **CUDA Version:** 12.8 (System)
* **Python Version:** 3.12
* **PyTorch Version:** 2.8.0+cu128

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with deep_gemm/include/deep_gemm/impls/smxx_clean_logits.cuh at line 17 and inspect how neg_inf is used by the fp8_mqa_logits kernel. Reproduce the NVRTC compilation failure in the stated CUDA 12.8 environment, then verify that the revised constant compiles and preserves the masking behavior without errors.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
compilers
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
72/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.