Lightning-AI / Lightning-AI/lightning-thunder
Create a parametrized benchmark for LitGPT SplitQKV+RoPE+Activation
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 1.5k
- Forks
- 121
- PR merge metrics
- No merged PRs in 30d
Description
🚀 Feature
Create a parametrized benchmark for LitGPT's parallel residual path (https://github.com/Lightning-AI/litgpt/blob/aaed893f858f350e22bb75df53b41f6bc0ff55ff/litgpt/model.py#L181).
The following two regions of the model code can be horizontally fused when parallel residual option is used:
- https://github.com/Lightning-AI/litgpt/blob/aaed893f858f350e22bb75df53b41f6bc0ff55ff/litgpt/model.py#L245-L262
- https://github.com/Lightning-AI/litgpt/blob/aaed893f858f350e22bb75df53b41f6bc0ff55ff/litgpt/model.py#L352
This benchmark should be parametrized similarly to test_litgpt_qkv_split_rope
https://github.com/Lightning-AI/lightning-thunder/blob/7d62ae1c7da8fb57e09da3e95ced734e34480400/thunder/benchmarks/targets.py#L572-L591
cc @crcrpar
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in thunder/benchmarks/targets.py with test_litgpt_qkv_split_rope, then inspect the two referenced regions in LitGPT's model.py. Add a parametrized benchmark for the parallel residual path covering SplitQKV, RoPE, and activation fusion; done means the benchmark runs across its parameters and measures the requested fused regions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- performance
- Issue type
- Feature
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 45/100