NVIDIA / NVIDIA/CUDALibrarySamples
[BUG] Cascaded algorithm correctness depends on compiler flag
@naveenaero is already working on this.
Since Nov 4, 2024.
- Dominant language
- Cuda
- Stars
- 2.5k
- Forks
- 478
- PR merge metrics
- No merged PRs in 30d
Description
Describe the bug
On branch-2.2, compiling the Cascaded algorithm with the -G flag changes the static shared memory alignment behavior, causing misalignment errors.
Steps/Code to reproduce bug
- Modify CMakeLists.txt:
- Comment out line CMakeLists.txt:33
set(CMAKE_CUDA_FLAGS_DEBUG "${CMAKE_CUDA_FLAGS_DEBUG};-g")
- Uncomment line CMakeLists.txt:32
set(CMAKE_CUDA_FLAGS_DEBUG "${CMAKE_CUDA_FLAGS_DEBUG};-G").
- Run tests/test_cascaded.cpp with the following test case:
TEST_CASE("comp/decomp cascaded-small-uint64", "[nvcomp][small]").
To compile successfully, I reduced the default_chunk_size in default_chunk_size from 4096 to 2048.
Expected behavior
The test should pass without errors.
Environment details (please complete the following information):
-
Environment location:
-
Ubuntu-22.04
-
Driver Version: 555.99
-
CUDA Version: 12.5
-
NVIDIA GeForce RTX 3080
-
Method of nvCOMP install: branch-2.2 source code
Additional context
After debugging, I explicitly declared alignment for shared memory allocation in the following files:
__shared__ __align__(sizeof(data_type)) uint8_t shmem[shmem_size];
__shared__ __align__(sizeof(data_type)) uint32_t chunk_metadata[max_chunk_metadata_size / sizeof(uint32_t)];
After making these changes, all tests in test_cascaded.cpp passed. I believe this dependency on compiler optimization for correctness is a bug.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.