Azure / Azure/cyclecloud-slurm

Compute nodes don't get terminated despite Slurm indicating idle~ status

Open
#267 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
84
Forks
56
Avg merge
3d 17h
Merged PRs (30d)
1

Description

CycleCloud version: 8.6.2-3276
Slurm version: 22.05.11

Autoscaling down after the job queue gets empty worked for me successfully numerous times until it didn't after full occupancy of the cluster lasting several days. All jobs were then killed, the compute nodes transitioned to `idle~`, but CycleCloud didn't deprovision the VMs. How can I investigate the cause of this behavior?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the interaction between CycleCloud 8.6.2-3276 and Slurm 22.05.11 when jobs are killed and compute nodes enter idle~ status. Reproduce the behavior after sustained full cluster occupancy and inspect the relevant logs to identify why deprovisioning is not triggered. Done means the cause is documented and node termination resumes after the queue becomes empty.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
cloud, hpc, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.