h2oai / h2oai/h2o4gpu

cuda9 volta DGX server slow, especially on ec2 instance

Open
#375 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
468
Forks
96
PR merge metrics
No merged PRs in 30d

Description

From nvidia testing:
```
Done for CUDA8 GBM and Kmeans_Image performance test, please reference below and attach. Thanks.
H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo GBM: GPU usage 37%~90%, please reference attach 1227_H2O4GPU_0.2.0_DGXServerPascal_CUDA8_NCCL_GBM.jpg.
H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo Kmeans_Image: GPU usage ~40%, please reference attach 1227_H2O4GPU_0.2.0_DGXServerPascal_CUDA8_NCCL_KmeansImage.jpg.

Summary:
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~20%.
- H2O4GPU_0.2.0_CUDA9_NoNCCL + DGX Server Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~20%.
- H2O4GPU_0.1.0_CUDA9_NCCL + DGX Server Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~20%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Station Volta with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~50%.
- H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~90%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta on AMI with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~10%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Pascal with demo Multi_GPU_H2O_GLM_SIMPLE: GPU usage ~40%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta with demo Kmeans_Image: GPU usage ~40%.
- H2O4GPU_0.2.0_CUDA9_NCCL + DGX Server Volta with demo GBM: GPU usage 20%~80%.
- H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo Kmeans_Image: GPU usage ~40%.
- H2O4GPU_0.2.0_CUDA8_NCCL + DGX Server Pascal with demo GBM: GPU usage 37%~90%.

```

Contributor guide

Open the contributing guide

Research direction

No source file or test is named. Start by reproducing Multi_GPU_H2O_GLM_SIMPLE, Kmeans_Image, and GBM across the listed CUDA 8/9, Pascal/Volta, DGX, and EC2 configurations, then compare GPU utilization. Done means identifying the configuration-dependent cause of low utilization and documenting a verified fix or narrowed reproduction.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.