babysor / babysor/MockingBird

GPU运算如何优化环境达到最高速率?

Open
#727 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
36.9k
Forks
5.2k
PR merge metrics
No merged PRs in 30d

Description

image
image
同样的环境 3090 size:66 0.59 steps/s
3060 size:16 0.65 steps/s
实在是搞不懂 而且CUDA11.6时 运算速率会随着模型增大而逐步变慢,3060一开始时1.1 steps/s 运行二十几个小时后逐步掉到0.65 steps/s
requirements.txt里有好几个支持的版本是在30系显卡有冲突或者不兼容的(也有可能是和win11不匹配),作者能否重新配置一下版本信息?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing requirements.txt alongside the reported CUDA 11.6, Windows 11, and GPU configurations, then try to reproduce the steps-per-second decline. Done means documenting compatible dependency versions and determining whether the reported slowdown or 30-series incompatibility is reproducible.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.