NVIDIA / NVIDIA/TransformerEngine
Reduce CPU overheads in te Sequential GroupedLinear Op
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 3.5k
- Forks
- 831
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 65
Description
Recently graph safe support has been added to te Sequential GroupedLinear Op https://github.com/NVIDIA/TransformerEngine/pull/2923.
But it suffers from CPU overheads. Nail the bottlenecks and fix them
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading PR #2923 and tracing the graph-safe Sequential GroupedLinear Op path. Profile it to identify the CPU bottlenecks, then verify that the overhead is reduced without regressions in the operation’s existing behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 38/100