Azure / Azure/cyclecloud-slurm

Latest CycleCloud (8.8.1) and Slurm clusters seems to have initialisation issues with GPU Nodes?

Open
#468 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
84
Forks
56
Avg merge
3d 17h
Merged PRs (30d)
1

Description

I have just installed CycleCloud 8.8.1, Slurm Template 4.0.4 and running Slurm 25.05.2 which comes with it all.

Normal nodes seem to be fine, but GPU Enabled nodes all throw up initialisation errors in the CC UI and show as failed (and pending/drained with sinfo) before they get terminated and another node is tried repeatedly. I have tried this with OOTB Alma 8, Alma 9 and my own Custom Alma 9 images for both scheduler and execute nodes.

Image

Image

Image

I have also seen these errors on another GPU node:

Image

ChatGPT seems to suggest that at least some of these are benign and that it is just the newer version of CC or the software versions in the latest HPC images that is being overly "picky" testing stuff, but even so, it still stops my GPU solving jobs from running.

Any ideas what is causing this and how to remedy?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.