Azure / Azure/cyclecloud-slurm
Compute nodes don't get terminated despite Slurm indicating idle~ status
- Dominant language
- Python
- Stars
- 84
- Forks
- 56
- Avg merge
- 3d 17h
- Merged PRs (30d)
- 1
Description
CycleCloud version: 8.6.2-3276
Slurm version: 22.05.11
Autoscaling down after the job queue gets empty worked for me successfully numerous times until it didn't after full occupancy of the cluster lasting several days. All jobs were then killed, the compute nodes transitioned to `idle~`, but CycleCloud didn't deprovision the VMs. How can I investigate the cause of this behavior?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the interaction between CycleCloud 8.6.2-3276 and Slurm 22.05.11 when jobs are killed and compute nodes enter idle~ status. Reproduce the behavior after sustained full cluster occupancy and inspect the relevant logs to identify why deprovisioning is not triggered. Done means the cause is documented and node termination resumes after the queue becomes empty.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- cloud, hpc, infrastructure
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100