AI-Hypercomputer / AI-Hypercomputer/xpk

NCCL is installed on cluster GPU hosts by default

Abierto
#1,014 3 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Python
Estrellas
193
Forks
94
Merge medio
24 min
PR fusionados (30 d)
1

Descripción

`NCCL` is installed on cluster GPU hosts by default (e.g on `a3-megagpu-8g` clusters) while this may not be the most desirable default behaviour.

User-containers (the '`gpu-image`' in `xpk`) typically bundle their own `NCCL` version that user code depends on and expects to use. However, in GKE we have found that this expectation is broken because `NCCL` binaries from GPU hosts are mounted into the containers and take precedence in the `LD_LIBRARY_PATH` by default. This could cause user code to break if there are discrepancies between the container and host NCCL versions.

It would be helpful to be able to easily disable `NCCL` install on GPU hosts by the `nccl-tcpxo-installer` pods. The manifest provided to `xpk` from `container-engine-accelerators` provides a way to do this by removing the `-install-nccl` flag ([nccl-tcpxo-installer.yaml#L88](https://github.com/GoogleCloudPlatform/container-engine-accelerators/blob/master/gpudirect-tcpxo/nccl-tcpxo-installer.yaml#L88))

With `xpk` cluster creation I was only able to do this by manually editing the deployment blueprints after being generated by xpk (and then manually deploying that). I was not able to get an end-to-end cluster deployment with `xpk` by, for example, repointing the hardcoded paths to the manifest ([e.g.](https://github.com/AI-Hypercomputer/xpk/blob/release-1.0/src/xpk/core/system_characteristics.py#L28)).

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.