NVIDIA / NVIDIA/cutlass

[BUG] Example 63 requires C++20 support

Open
#3,011 7 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

? - Needs Triage bug CUTLASS C++ inactive-90d
Dominant language
C++
Stars
10.5k
Forks
2.1k
Avg merge
3d 11h
Merged PRs (30d)
7

Description

Which component has the problem?

CUTLASS C++

Bug Report

Describe the bug
Compilation using NVHPC 25.9 fails with the following error:

/e/project1/jureap5/reuter8/raps/develop/external/cutlass/examples/63_hopper_gemm_with_weight_prefetch/collective/sm90_mma_tma_gmma_ss_warpspeciali
zed_with_prefetch.hpp: In instantiation of ‘constexpr const auto cutlass::gemm::collective::CollectiveMma<cutlass::gemm::MainloopSm90TmaGmmaWarpSpe
cializedWithPrefetch<13, cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> >, cutlass::gemm::KernelTmaWarpSpecializedFP8FastAccumWithPrefetchAndSplitD
MA>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cutlass::float_e4m3_t, cute::tuple<long int, cute::C<1>, long int>, cutlass::float_e5m2_
t, cute::tuple<long int, cute::C<1>, long int>, cute::TiledMMA<cute::MMA_Atom<cute::SM90::GMMA::MMA_64x64x32_F32E4M3E5M2_SS_TN<cute::SM90::GMMA::Sc
aleIn::One, cute::SM90::GMMA::ScaleIn::One> >, cute::Layout<cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> > >, cute::tuple<cute::Underscore, cute:
:Underscore, cute::Underscore> >, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<3, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute:
:tuple<cute::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, void, cute::identity, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute
::Swizzle<3, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, v
oid, cute::identity>::prefetch_smem_size’:
/e/project1/jureap5/reuter8/raps/develop/external/cutlass/examples/63_hopper_gemm_with_weight_prefetch/collective/sm90_mma_tma_gmma_ss_warpspeciali
zed_with_prefetch.hpp:164:30:   required from ‘struct cutlass::gemm::collective::CollectiveMma<cutlass::gemm::MainloopSm90TmaGmmaWarpSpecializedWit
hPrefetch<13, cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> >, cutlass::gemm::KernelTmaWarpSpecializedFP8FastAccumWithPrefetchAndSplitDMA>, cute::
tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cutlass::float_e4m3_t, cute::tuple<long int, cute::C<1>, long int>, cutlass::float_e5m2_t, cute::tu
ple<long int, cute::C<1>, long int>, cute::TiledMMA<cute::MMA_Atom<cute::SM90::GMMA::MMA_64x64x32_F32E4M3E5M2_SS_TN<cute::SM90::GMMA::ScaleIn::One,
 cute::SM90::GMMA::ScaleIn::One> >, cute::Layout<cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> > >, cute::tuple<cute::Underscore, cute::Underscore
, cute::Underscore> >, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<3, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute
::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, void, cute::identity, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<3
, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, void, cute::
identity>’
  164 |   static constexpr auto prefetch_smem_size = cute::cosize_v<PrefetchSmemLayoutA>;
      |                              ^~~~~~~~~~~~~~~~~~
/e/project1/jureap5/reuter8/raps/develop/external/cutlass/include/cutlass/gemm/device/gemm_universal_adapter.h:630:137:   recursively required by s
ubstitution of ‘template<class GemmKernel> struct cutlass::gemm::detail::IsCutlass3GemmKernel<GemmKernel, std::void_t<typename GemmKernel::ProblemS
hape> > [with GemmKernel = cutlass::gemm::kernel::GemmUniversal<cute::tuple<int, int, int, int>, cutlass::gemm::collective::CollectiveMma<cutlass::
gemm::MainloopSm90TmaGmmaWarpSpecializedWithPrefetch<13, cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> >, cutlass::gemm::KernelTmaWarpSpecializedF
P8FastAccumWithPrefetchAndSplitDMA>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cutlass::float_e4m3_t, cute::tuple<long int, cute::C<1>,
 long int>, cutlass::float_e5m2_t, cute::tuple<long int, cute::C<1>, long int>, cute::TiledMMA<cute::MMA_Atom<cute::SM90::GMMA::MMA_64x64x32_F32E4M
3E5M2_SS_TN<cute::SM90::GMMA::ScaleIn::One, cute::SM90::GMMA::ScaleIn::One> >, cute::Layout<cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> > >, cut
e::tuple<cute::Underscore, cute::Underscore, cute::Underscore> >, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<3, 4, 3>, cute::smem_ptr_
flag_bits<8>, cute::Layout<cute::tuple<cute::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, void, cute::identity, cute::SM90_TMA
_LOAD, cute::ComposedLayout<cute::Swizzle<3, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<8>, cute::C<128> >, cute::tuple<c
ute::C<128>, cute::C<1> > > >, void, cute::identity>, cutlass::epilogue::collective::CollectiveEpilogue<cutlass::epilogue::Sm90TmaWarpSpecialized<1
, 1, 32, false, false>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cute::tuple<cute::C<64>, cute::C<64> >, cutlass::float_e4m3_t, cute::
tuple<cute::C<1>, long int, long int>, cutlass::float_e4m3_t, cute::tuple<cute::C<1>, long int, long int>, cutlass::epilogue::fusion::FusionCallbac
ks<cutlass::epilogue::Sm90TmaWarpSpecialized<1, 1, 32, false, false>, cutlass::epilogue::fusion::LinearCombination<cutlass::float_e4m3_t, float, cu
tlass::float_e4m3_t, float, cutlass::FloatRoundStyle::round_to_nearest>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cute::tuple<cute::C<
64>, cute::C<64> > >, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<2, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute:
:C<64>, cute::C<8> >, cute::tuple<cute::C<1>, cute::C<64> > > >, cute::AutoVectorizingCopyWithAssumedAlignment<128>, cute::SM90_TMA_STORE, cute::Co
mposedLayout<cute::Swizzle<2, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<64>, cute::C<8> >, cute::tuple<cute::C<1>, cute:
:C<64> > > >, cute::AutoVectorizingCopyWithAssumedAlignment<128>, cute::Copy_Atom<cute::SM90_U32x4_STSM_N, cutlass::half_t>, void>, void>]’
  630 | class GemmUniversalAdapter<
      |

/e/project1/jureap5/reuter8/raps/develop/external/cutlass/include/cutlass/gemm/device/gemm_universal_adapter.h:630:137:   required by substitution
of ‘template<class GemmKernel_> class cutlass::gemm::device::GemmUniversalAdapter<GemmKernel_, typename std::enable_if<(! cutlass::gemm::detail::IsCutlass3GemmKernel<typename cutlass::GetUnderlyingKernel<T>::type>::value), void>::type> [with GemmKernel_ = cutlass::gemm::kernel::GemmUniversal<cute::tuple<int, int, int, int>, cutlass::gemm::collective::CollectiveMma<cutlass::gemm::MainloopSm90TmaGmmaWarpSpecializedWithPrefetch<13, cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> >, cutlass::gemm::KernelTmaWarpSpecializedFP8FastAccumWithPrefetchAndSplitDMA>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cutlass::float_e4m3_t, cute::tuple<long int, cute::C<1>, long int>, cutlass::float_e5m2_t, cute::tuple<long int, cute::C<1>, long int>, cute::TiledMMA<cute::MMA_Atom<cute::SM90::GMMA::MMA_64x64x32_F32E4M3E5M2_SS_TN<cute::SM90::GMMA::ScaleIn::One, cute::SM90::GMMA::ScaleIn::One> >, cute::Layout<cute::tuple<cute::C<1>, cute::C<1>, cute::C<1> > >, cute::tuple<cute::Underscore, cute::Underscore, cute::Underscore> >, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<3, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, void, cute::identity, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<3, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<8>, cute::C<128> >, cute::tuple<cute::C<128>, cute::C<1> > > >, void, cute::identity>, cutlass::epilogue::collective::CollectiveEpilogue<cutlass::epilogue::Sm90TmaWarpSpecialized<1, 1, 32, false, false>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cute::tuple<cute::C<64>, cute::C<64> >, cutlass::float_e4m3_t, cute::tuple<cute::C<1>, long int, long int>, cutlass::float_e4m3_t, cute::tuple<cute::C<1>, long int, long int>, cutlass::epilogue::fusion::FusionCallbacks<cutlass::epilogue::Sm90TmaWarpSpecialized<1, 1, 32, false, false>, cutlass::epilogue::fusion::LinearCombination<cutlass::float_e4m3_t, float, cutlass::float_e4m3_t, float, cutlass::FloatRoundStyle::round_to_nearest>, cute::tuple<cute::C<64>, cute::C<64>, cute::C<128> >, cute::tuple<cute::C<64>, cute::C<64> > >, cute::SM90_TMA_LOAD, cute::ComposedLayout<cute::Swizzle<2, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<64>, cute::C<8> >, cute::tuple<cute::C<1>, cute::C<64> > > >, cute::AutoVectorizingCopyWithAssumedAlignment<128>, cute::SM90_TMA_STORE, cute::ComposedLayout<cute::Swizzle<2, 4, 3>, cute::smem_ptr_flag_bits<8>, cute::Layout<cute::tuple<cute::C<64>, cute::C<8> >, cute::tuple<cute::C<1>, cute::C<64> > > >, cute::AutoVectorizingCopyWithAssumedAlignment<128>, cute::Copy_Atom<cute::SM90_U32x4_STSM_N, cutlass::half_t>, void>, void>]’
/e/project1/jureap5/reuter8/raps/develop/external/cutlass/examples/63_hopper_gemm_with_weight_prefetch/63_hopper_gemm_with_weight_prefetch.cu:199:0:   required from here
  199 | using EpilogueOutputOp  = typename Gemm::EpilogueOutputOp;
/e/project1/jureap5/reuter8/raps/develop/external/cutlass/examples/63_hopper_gemm_with_weight_prefetch/collective/sm90_mma_tma_gmma_ss_warpspecialized_with_prefetch.hpp:159:63: error: non-type template parameters of class type only available with ‘-std=c++20’ or ‘-std=gnu++20’
  159 |   using PrefetchSmemLayoutA = decltype(make_layout(make_shape(
      |                                                     ~~~~~~~~~~^
/e/project1/jureap5/reuter8/raps/develop/external/cutlass/examples/63_hopper_gemm_with_weight_prefetch/collective/sm90_mma_tma_gmma_ss_warpspecialized_with_prefetch.hpp:159:63: error: non-type template parameters of class type only available with ‘-std=c++20’ or ‘-std=gnu++20’
gmake[2]: *** [examples/63_hopper_gemm_with_weight_prefetch/CMakeFiles/63_hopper_gemm_with_weight_prefetch.dir/build.make:80: examples/63_hopper_gemm_with_weight_prefetch/CMakeFiles/63_hopper_gemm_with_weight_prefetch.dir/63_hopper_gemm_with_weight_prefetch.cu.o] Error 1
gmake[1]: *** [CMakeFiles/Makefile2:61836: examples/63_hopper_gemm_with_weight_prefetch/CMakeFiles/63_hopper_gemm_with_weight_prefetch.dir/all] Error 2
gmake[1]: *** Waiting for unfinished jobs....

Steps/Code to reproduce bug

Compile Cutlass v4.3.5 using the following commands:

export FC=nvfortran
export CC=nvc
export CXX=nvc++
cmake -DCMAKE_INSTALL_PREFIX=$installdir \
          -DCMAKE_C_COMPILER=$CC \
          -DCMAKE_CXX_COMPILER=$CXX \
          -DCMAKE_C_COMPILER_FLAGS="-O2 -fPIC" \
          -DCMAKE_CXX_COMPILER_FLAGS="-O2 -fPIC" \
          -DCUTLASS_NVCC_ARCHS="90" \
      -S $srcdir -B $builddir

cmake --build $builddir -j ${NCPUS:-32}

Expected behavior
Successful build. I managed to work around it by hacking the build flags to use -std=c++20 in $builddir/examples/63_hopper_gemm_with_weight_prefetch/CMakeFiles/63_hopper_gemm_with_weight_prefetch.dir/flags.make

Environment details (please complete the following information):
JSC JUPITER Supercomputer. NVHPC 25.9.

Additional context
Add any other context about the problem here.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with examples/63_hopper_gemm_with_weight_prefetch/63_hopper_gemm_with_weight_prefetch.cu and collective/sm90_mma_tma_gmma_ss_warpspecialized_with_prefetch.hpp, then inspect the generated flags.make from the reported CMake build. Reproduce with NVHPC 25.9 and verify that the example builds without manually adding -std=c++20; the C++20 requirement should be handled by the project configuration.

Written by the indexing model from the issue text.

Assessment

Tech stack
cmake, cpp
Domain
build-system
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.