ClusterLabs / ClusterLabs/resource-agents

azure-events-az: Node Health Attribute (-1000000) Not Automatically Reset After Critical Failure in Resource-Agent Update

Open
#2,025 7 comments 0 reactions 0 assignees View on GitHub
Dominant language
Shell
Stars
519
Forks
608
Avg merge
6d 1h
Merged PRs (30d)
7

Description

We have observed that with the recent update to the resource-agent configuration, as outlined in [Cluster Lab resource-agents](https://github.com/ClusterLabs/resource-agents/blob/90f9f1c43f3b1cb0611f31ab82142614f5cb7bb7/heartbeat/azure-events-az.in#L535), once a node is marked with a value of `-1000000`—indicating a critical failure or an unhealthy state—the attribute remains unchanged until it is manually modified. Consequently, cluster services will not restart until this manual intervention occurs.

Could you clarify the rationale behind why the value is not automatically reset to `0`?

Contributor guide

No contributing guide indexed for this repository

Research direction

Read heartbeat/azure-events-az.in at the referenced line 535 and review the linked resource-agent configuration update. Determine why a node attribute of -1000000 remains unchanged after a critical failure, and document whether automatic reset to 0 is expected; the issue does not define a code change or test for completion.

Written by the indexing model from the issue text.

Assessment

Tech stack
azure, shell
Domain
cloud, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.