support custom triton sdpa kernel in CUDA backend
未关闭
@Gasoonjia 已经在做这个了。
开始于 2025年11月13日。
- 主要语言
- Python
- 星标
- 5k
- 派生
- 1.2k
- 平均合并
- 2 天 10 小时
- 30 天内合并 PR
- 581
描述
🚀 The feature, motivation and pitch
Currnently one of the main gap for cuda backend is we don't support sdpa kernel in one step, but need to decompose it which introduces extra perf latency.
We should have a single triton sdpa kernel for CUDA backend.
Alternatives
No response
Additional context
No response
RFC (Optional)
No response
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
评估
这个 Issue 还没有评估数据。