microsoft / microsoft/TRELLIS.2

RuntimeError: cuDNN error: CUDNN_STATUS_NOT_INITIALIZED

Open
#154 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
11.3k
Forks
1.4k
PR merge metrics
No merged PRs in 30d

Description

The Conda environment is shown below.
name: trellis2
channels:

  • defaults
    dependencies:
  • _libgcc_mutex=0.1=main
  • _openmp_mutex=5.1=1_gnu
  • bzip2=1.0.8=h5eee18b_6
  • ca-certificates=2026.3.19=h06a4308_0
  • ld_impl_linux-64=2.44=h9e0c5a2_3
  • libexpat=2.7.5=h7354ed3_0
  • libffi=3.4.4=h6a678d5_1
  • libgcc=15.2.0=h69a1729_7
  • libgcc-ng=15.2.0=h166f726_7
  • libgomp=15.2.0=h4751f2c_7
  • libnsl=2.0.0=h5eee18b_0
  • libstdcxx=15.2.0=h39759b7_7
  • libstdcxx-ng=15.2.0=hc03a8fd_7
  • libuuid=1.41.5=h5eee18b_0
  • libxcb=1.17.0=h9b100fa_0
  • libzlib=1.3.1=h47b2149_1
  • ncurses=6.5=h7934f7d_0
  • openssl=3.5.6=h1b28b03_0
  • packaging=26.0=py310h06a4308_0
  • pip=26.0.1=pyhc872135_1
  • pthread-stubs=0.3=h0ce48e5_1
  • python=3.10.20=h741d88c_0
  • readline=8.3=hc2a1206_0
  • sqlite=3.51.2=h3e8d24a_0
  • tk=8.6.15=h54e0aa7_0
  • wheel=0.46.3=py310h06a4308_0
  • xorg-libx11=1.8.12=h9b100fa_1
  • xorg-libxau=1.0.12=h9b100fa_0
  • xorg-libxdmcp=1.1.5=h9b100fa_0
  • xorg-xorgproto=2024.1=h5eee18b_1
  • xz=5.8.2=h448239c_0
  • zlib=1.3.1=h47b2149_1
  • pip:
    • absl-py==2.4.0
    • aiofiles==24.1.0
    • annotated-doc==0.0.4
    • annotated-types==0.7.0
    • anyio==4.13.0
    • brotli==1.2.0
    • certifi==2026.4.22
    • click==8.3.3
    • cuda-bindings==13.2.0
    • cuda-pathfinder==1.5.3
    • cuda-toolkit==13.0.2
    • cumesh==0.0.1
    • easydict==1.13
    • einops==0.8.2
    • exceptiongroup==1.3.1
    • fastapi==0.136.1
    • ffmpy==1.0.0
    • filelock==3.29.0
    • flash-attn==2.7.3
    • flex-gemm==1.0.0
    • fsspec==2026.3.0
    • glcontext==3.0.0
    • gradio==6.0.1
    • gradio-client==2.0.0
    • groovy==0.1.2
    • grpcio==1.80.0
    • h11==0.16.0
    • hf-xet==1.4.3
    • httpcore==1.0.9
    • httpx==0.28.1
    • huggingface-hub==1.11.0
    • idna==3.13
    • imageio==2.37.3
    • imageio-ffmpeg==0.6.0
    • jinja2==3.1.6
    • kornia==0.8.2
    • kornia-rs==0.1.10
    • lpips==0.1.4
    • markdown==3.10.2
    • markdown-it-py==4.0.0
    • markupsafe==3.0.3
    • mdurl==0.1.2
    • moderngl==5.12.0
    • mpmath==1.3.0
    • networkx==3.4.2
    • ninja==1.13.0
    • numpy==2.2.6
    • nvdiffrast==0.4.0
    • nvdiffrec-render==0.0.0
    • nvidia-cublas==13.1.0.3
    • nvidia-cublas-cu12==12.4.5.8
    • nvidia-cuda-cupti==13.0.85
    • nvidia-cuda-cupti-cu12==12.4.127
    • nvidia-cuda-nvrtc==13.0.88
    • nvidia-cuda-nvrtc-cu12==12.4.127
    • nvidia-cuda-runtime==13.0.96
    • nvidia-cuda-runtime-cu12==12.4.127
    • nvidia-cudnn-cu12==9.1.0.70
    • nvidia-cudnn-cu13==9.21.1.3
    • nvidia-cufft==12.0.0.61
    • nvidia-cufft-cu12==11.2.1.3
    • nvidia-cufile==1.15.1.6
    • nvidia-curand==10.4.0.35
    • nvidia-curand-cu12==10.3.5.147
    • nvidia-cusolver==12.0.4.66
    • nvidia-cusolver-cu12==11.6.1.9
    • nvidia-cusparse==12.6.3.3
    • nvidia-cusparse-cu12==12.3.1.170
    • nvidia-cusparselt-cu12==0.6.2
    • nvidia-cusparselt-cu13==0.8.0
    • nvidia-nccl-cu12==2.21.5
    • nvidia-nccl-cu13==2.28.9
    • nvidia-nvjitlink==13.0.88
    • nvidia-nvjitlink-cu12==12.4.127
    • nvidia-nvshmem-cu13==3.4.5
    • nvidia-nvtx==13.0.85
    • nvidia-nvtx-cu12==12.4.127
    • o-voxel==0.0.1
    • opencv-python-headless==4.13.0.92
    • orjson==3.11.8
    • pandas==2.3.3
    • pillow==12.1.1
    • pillow-simd==9.5.0.post2
    • plyfile==1.1.3
    • protobuf==7.34.1
    • psutil==7.2.2
    • pydantic==2.12.4
    • pydantic-core==2.41.5
    • pydub==0.25.1
    • pygments==2.20.0
    • python-dateutil==2.9.0.post0
    • python-multipart==0.0.26
    • pytz==2026.1.post1
    • pyyaml==6.0.3
    • regex==2026.4.4
    • rich==15.0.0
    • safehttpx==0.1.7
    • safetensors==0.7.0
    • scipy==1.15.3
    • semantic-version==2.10.0
    • setuptools==81.0.0
    • shellingham==1.5.4
    • six==1.17.0
    • starlette==0.52.1
    • sympy==1.13.1
    • tensorboard==2.20.0
    • tensorboard-data-server==0.7.2
    • timm==1.0.26
    • tokenizers==0.22.2
    • tomlkit==0.13.3
    • torch==2.6.0+cu124
    • torchaudio==2.6.0+cu124
    • torchvision==0.21.0+cu124
    • tqdm==4.67.3
    • transformers==5.6.2
    • trimesh==4.12.0
    • triton==3.2.0
    • typer==0.24.2
    • typing-extensions==4.15.0
    • typing-inspection==0.4.2
    • tzdata==2026.1
    • utils3d==0.0.2
    • uvicorn==0.46.0
    • werkzeug==3.1.8
    • zstandard==0.25.0

Run the command: cat /usr/include/cudnn_version.h | grep CUDNN_MAJOR -A 2
The following information is returned.
#define CUDNN_MAJOR 9
#define CUDNN_MINOR 0
#define CUDNN_PATCHLEVEL 0

#define CUDNN_VERSION (CUDNN_MAJOR * 10000 + CUDNN_MINOR * 100 + CUDNN_PATCHLEVEL)

/* cannot use constexpr here since this is a C-only file */

After checking relevant information, this issue appears to be caused by a cuDNN version mismatch. Could you please advise on feasible solutions to resolve this problem?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by comparing the cuDNN version reported by /usr/include/cudnn_version.h with the installed nvidia-cudnn and torch==2.6.0+cu124 packages listed in the environment. Reproduce the CUDNN_STATUS_NOT_INITIALIZED error in this Conda environment and verify that a compatible dependency configuration resolves it.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.