support custom triton sdpa kernel in CUDA backend
Đang mở
@Gasoonjia đang làm issue này rồi.
Từ ngày 13/11/2025.
- Ngôn ngữ chính
- Python
- Star
- 5k
- Fork
- 1.2k
- Merge trung bình
- 2 ngày 10 giờ
- Pull request đã merge (30 ngày)
- 581
Mô tả
🚀 The feature, motivation and pitch
Currnently one of the main gap for cuda backend is we don't support sdpa kernel in one step, but need to decompose it which introduces extra perf latency.
We should have a single triton sdpa kernel for CUDA backend.
Alternatives
No response
Additional context
No response
RFC (Optional)
No response
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Đánh giá
Issue này chưa được đánh giá.