NVIDIA / NVIDIA/cutlass

[QST] which is optimised way to iterate over the conv2d filter tensor

Open
#2,146 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

? - Needs Triage inactive-30d inactive-90d question
Dominant language
C++
Stars
10.5k
Forks
2.1k
Avg merge
3d 11h
Merged PRs (30d)
7

Description

What is your question?
Hi, I want to iterate and perform some operations over the elements of every channel of each filter's filter tensor of a conv2dfprop operation.

According to https://github.com/NVIDIA/cutlass/blob/main/media/docs/implicit_gemm_convolution.md, the filters are grouped inside the tensor in shapes K, R, S, C.

I want to perform an adding operation of every element of each channel from each filter. To perform that, I need to iterate through each K filter and, inside each C channel, group its R*S elements per channel and perform the total addition of them.

As I aim to perform this inside the kernel (https://github.com/NVIDIA/cutlass/blob/24f991e87930e1159f1f5a47e329d43bcfbd76b9/include/cutlass/conv/kernel/implicit_gemm_convolution.h#L281) and with the threads of the GPU, I wanted to parallelize this iteration.

From my point of view, the best way is for each thread of the GPU (in my case, a GPU Tesla v100 with 128 threads) to add a channel's elements. So, each thread needs to process an addition of RS elements. So if we have KC channels, I need different K*C threads so no thread iterates over others' data. If we don't have enough threads, one thread could perform 2 channels. The code should be similar to the following:

// Shape of Filters Tensor
        long int K = params.problem_size.K;  // Number of filters
        long int C = params.problem_size.C;  // Channels per filter
        long int R = params.problem_size.R;  // Height of the filter 
        long int S = params.problem_size.S;  // Width of the filter
        long int HW = R * S;                 // Number of elements per channel of a filter
        long int total_channels = K * C;

        // Ensure each thread processes multiple channels if there are more channels than threads
        for (long int channel_idx = thread_idx; channel_idx < total_channels; channel_idx += total_threads) {

          // Compute (filter_id, channel_id) for the given channel index
          long int n = channel_idx / C;  // Filter index
          long int c = channel_idx % C;  // Channel index within the filter

          // starting index of each thread within the tensor
          long int base_index = n * (C * HW) + c;


          // iteration over R*S elements per channel 
         float sum = 0.0
          for (long int i = 0; i < HW; i++) {
            // the elements of each channel are stored in memory i*C elements apart
             sum  += params.ptr_B[base_index + i * C]);
          }

Am I doing it properly, or am I iterating over the tensor in a suboptimal way so I am downgrading GPU performance?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with media/docs/implicit_gemm_convolution.md and the linked section of include/cutlass/conv/kernel/implicit_gemm_convolution.h. Compare the proposed K/C/R/S indexing and thread mapping with the convolution kernel's existing tensor layout and launch structure. Done would require a confirmed, benchmarked iteration strategy or a clearly scoped follow-up change.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
hpc, performance
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.