NVIDIA / NVIDIA/cutlass

[QST] Support fp8 gemm with 128x1 LHS scaling and 1x128 RHS scaling

Open
#2,280 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

? - Needs Triage inactive-30d inactive-90d question
Dominant language
C++
Stars
10.5k
Forks
2.1k
Avg merge
3d 11h
Merged PRs (30d)
7

Description

In 67_hopper_fp8_warp_specialized_gemm_with_groupwise_scaling, if the params are changed like this

constexpr int ScaleGranularityM = 128; 
constexpr int ScaleGranularityN = 128;
constexpr int ScaleGranularityK = 1;

dose it support fp8 gemm with 128x1 LHS scaling and 1x128 RHS scaling?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with examples/67_hopper_fp8_warp_specialized_gemm_with_blockwise_scaling/67_hopper_fp8_warp_specialized_gemm_with_groupwise_scaling.cu and inspect the ScaleGranularityM, ScaleGranularityN, and ScaleGranularityK parameters. Determine whether the stated 128x1 and 1x128 scaling configuration is supported, and verify the conclusion with the example's available build or test path.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
hpc, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.