cuda9 volta DGX server slow, especially on ec2 instance
- Dominant language
- C++
- Stars
- 468
- Forks
- 96
- PR merge metrics
- No merged PRs in 30d
Description
From nvidia testing:
```
Done for CUDA8 GBM and Kmeans_Image performance test, please reference below and attach. Thanks.
H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo GBM: GPU usage 37%~90%, please reference attach 1227_H2O4GPU_0.2.0_DGXServerPascal_CUDA8_NCCL_GBM.jpg.
H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo Kmeans_Image: GPU usage ~40%, please reference attach 1227_H2O4GPU_0.2.0_DGXServerPascal_CUDA8_NCCL_KmeansImage.jpg.
Summary:
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~20%.
- H2O4GPU_0.2.0_CUDA9_NoNCCL + DGX Server Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~20%.
- H2O4GPU_0.1.0_CUDA9_NCCL + DGX Server Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~20%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Station Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~50%.
- H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~90%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta on AMI with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~10%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Pascal with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~40%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta with demo Kmeans_Image: GPU usage ~40%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta with demo GBM: GPU usage 20%~80%.
- H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo Kmeans_Image: GPU usage ~40%.
- H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo GBM: GPU usage 37%~90%.
```
Contributor guide
Research direction
No source file or test is named. Start by reproducing Multi_GPU_H2O_GLM_SIMPLE, Kmeans_Image, and GBM across the listed CUDA 8/9, Pascal/Volta, DGX, and EC2 configurations, then compare GPU utilization. Done means identifying the configuration-dependent cause of low utilization and documenting a verified fix or narrowed reproduction.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100