kohya-ss / kohya-ss/sd-scripts
CUDNN 9.0
- Dominant language
- Python
- Stars
- 7.2k
- Forks
- 1.2k
- Avg merge
- 11m
- Merged PRs (30d)
- 2
Description
https://docs.nvidia.com/deeplearning/cudnn/release-notes.html#cudnn-9-0-0
Is it possible to add Cudnn 9.0
The cuDNN backend API uses less memory for many execution plans which should be beneficial for users who cache execution plans.
FP16 and BF16 fused flash attention engine performance has been significantly improved for NVIDIA GPUs:
Speed-up of up to 50% over cuDNN 8.9.7 on Hopper GPUs.
Speed-up of up to 100% over cuDNN 8.9.7 on Ampere GPUs.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the linked NVIDIA cuDNN 9.0 release notes, then inspect the repository's dependency and setup configuration to find where cuDNN compatibility is selected. Determine the required update and verify it against the project's existing checks or supported installation path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100