pytorch / pytorch/executorch

support custom triton sdpa kernel in CUDA backend

未关闭
#15,744 2 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

@Gasoonjia 已经在做这个了。

开始于 2025年11月13日。

主要语言
Python
星标
5k
派生
1.2k
平均合并
2 天 10 小时
30 天内合并 PR
581

描述

🚀 The feature, motivation and pitch

Currnently one of the main gap for cuda backend is we don't support sdpa kernel in one step, but need to decompose it which introduces extra perf latency.
We should have a single triton sdpa kernel for CUDA backend.

Alternatives

No response

Additional context

No response

RFC (Optional)

No response

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。