NVIDIA / NVIDIA/cutlass

example/58_ada_fp8_gemm

Open
#2,149 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

? - Needs Triage inactive-30d inactive-90d question
Dominant language
C++
Stars
10.5k
Forks
2.1k
Avg merge
3d 11h
Merged PRs (30d)
7

Description

What is your question?

using ElementA = cutlass::float_e4m3_t;
using ElementB = cutlass::float_e4m3_t;
using ElementOutput = cutlass::float_e4m3_t;
using ElementAuxOutput = ElementOutput;
using ElementAccumulator = float; -------------------->>cutlass::half_t

Is it feasible for me to convert the FP32 accumulator into the FP16 accumulator?, Why does this run with an error? I need to write the instance myself

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with example/58_ada_fp8_gemm and reproduce the error using the shown ElementAccumulator aliases. Inspect the example's CUTLASS configuration to determine whether a cutlass::half_t accumulator is supported for this FP8 GEMM. Done means confirming a supported configuration or documenting the required instance and the cause of the error.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.