AI-Hypercomputer / AI-Hypercomputer/xpk

NCCL is installed on cluster GPU hosts by default

Ouverte
#1,014 3 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Python
Étoiles
193
Forks
94
Merge moyen
24 min
PR mergées (30 j)
1

Description

`NCCL` is installed on cluster GPU hosts by default (e.g on `a3-megagpu-8g` clusters) while this may not be the most desirable default behaviour.

User-containers (the '`gpu-image`' in `xpk`) typically bundle their own `NCCL` version that user code depends on and expects to use. However, in GKE we have found that this expectation is broken because `NCCL` binaries from GPU hosts are mounted into the containers and take precedence in the `LD_LIBRARY_PATH` by default. This could cause user code to break if there are discrepancies between the container and host NCCL versions.

It would be helpful to be able to easily disable `NCCL` install on GPU hosts by the `nccl-tcpxo-installer` pods. The manifest provided to `xpk` from `container-engine-accelerators` provides a way to do this by removing the `-install-nccl` flag ([nccl-tcpxo-installer.yaml#L88](https://github.com/GoogleCloudPlatform/container-engine-accelerators/blob/master/gpudirect-tcpxo/nccl-tcpxo-installer.yaml#L88))

With `xpk` cluster creation I was only able to do this by manually editing the deployment blueprints after being generated by xpk (and then manually deploying that). I was not able to get an end-to-end cluster deployment with `xpk` by, for example, repointing the hardcoded paths to the manifest ([e.g.](https://github.com/AI-Hypercomputer/xpk/blob/release-1.0/src/xpk/core/system_characteristics.py#L28)).

Guide de contribution

Ouvrir le guide de contribution

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.