[QST] How to define a new custom kernel
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 10.5k
- Forks
- 2.1k
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 7
Description
What is your question?
Hi, I want to define a new convolution2d kernel Fprop based on the bare convolution kernel.
My idea is to modify the weights or filters before the convolution operation.
I would need a little help to know or understand how Cutlass allows the definition of a new custom kernel and how it is executed later in the GPU (my GPU is a Tesla v100).
I am a little confused by looking at examples like the number 9 on how I must perform the modification of the weights or filters because it seems only to execute templates to other templates like default convolution but no operation is performed.
Any help or link to a clarifier example on how to program my kernel and later use it or a base explanation would be appreciated.
Thank you.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with examples/09_turing_tensorop_conv2dfprop and include/cutlass/conv/kernel/default_conv2d.h, then trace how the template composition leads to execution on the GPU. Done would be a clear explanation or example showing where custom weight or filter modification belongs and how the resulting kernel is used.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 18/100