kvcache-ai / kvcache-ai/ktransformers

How are you doing sparse attention for rtx blackwell? sm120

Open
#1,680 2 comments 0 reactions 2 assignees Claimed by @ovowei View on GitHub
bug
Dominant language
Python
Stars
19.5k
Forks
1.6k
Avg merge
19h 32m
Merged PRs (30d)
27

Description

### Reminder

- [x] I have read the above rules and searched the existing issues.

### System Info

this repo on HF Kebob/DeepSeek-V3.2-CPU-NUMA2-AMXINT4
claims to run this on RTX PRO 6000 sm120 and ktransformers
but flashmla the only sparse attention backend im aware of doesn't support sm120??
and sglang instructions for ds32 include installing flashmla...?

### Reproduction

```text
Put your message here.
```

### Others

_No response_

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.