NVIDIA / NVIDIA/TensorRT-LLM

How to use block_sparse in tensorrt_llm.layers.attention.Attention

Open
#6,378 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
14.7k
Forks
2.8k
Avg merge
2d 23h
Merged PRs (30d)
489

Description

Hi there,

I want to evaluate the performance of block_sparse attention (should be in tensorrt_llm.layers.attention.Attention). How is it implemented? Is there any tutorial for it?
I also noticed that tensorrt_llm.functional.gpt_attention has parameters for block_sparse, but I did not find anywhere that has an example for it.

Thanks!

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reading tensorrt_llm.layers.attention.Attention and tensorrt_llm.functional.gpt_attention, focusing on their block_sparse parameters and how they are used. Document the implementation, usage steps, and a tutorial or example that lets users evaluate block-sparse attention performance.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
documentation, machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.